A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
AuthorsSietse Schelpe
Resources
Instead of making the model smarter, this system remembers verified answers so it can return them instantly, exactly, and forever.
Key results
Accuracy across 180 fresh instances spanning nine problem families
Verified answers are produced without token generation
Full reuse completes within the reported 6–23 ms range
Wrong-item selection rate on a 4,500-item verified store
Token capacity measured on a single 46 GB GPU
What the paper found
An industry report from Sietse Schelpe of Corbenic AI presents Galahad, a proprietary system whose Merlin integrity layer stores solutions only after answer-key-free verification, while keeping the underlying language model frozen. Instead of regenerating answers, Galahad stores verified parameterized methods and executes them deterministically on new instances. Across 180 fresh instances from nine algorithmic and mathematical families, frozen Gemma-4-12B achieved 100% accuracy at 0 generation tokens, and the same result held for Qwen3-14B, DeepSeek-Coder-V2-Lite, and Phi-4, spanning dense and mixture-of-experts architectures. Reuse completed in a measured 6–23 ms, while consistency-gated reasoning achieved 88/88 accepted reuses and transferred methods correctly in 77/80 cases. The report also argues that exact addressing is essential: on a 4,500-item verified store, approximate similarity retrieval selected the wrong item 94.3% of the time, whereas exact addressing produced zero collisions. Galahad further maintained a movable 6,000,000-token context on one 46 GB GPU, compared with vLLM’s 30,399-token ceiling and SGLang’s silent truncation beyond roughly 32,000 tokens. The authors carefully scope the claim: Anthropic’s frontier models and Google DeepMind’s Gemma reporting still show superior cold-start reasoning, but on previously solved and verified workloads, this frozen-model approach trades repeated generation for bit-exact, auditable execution at near-zero marginal cost.
Original abstract
Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning nine problem families, four architectures from four vendors - dense and mixture-of-experts - each score 180/180 at zero generation tokens per answer: execution-bound capability decoupled from parameter scaling. A negative control attributes the capability fully to the memory: emptied, it solves nothing. The same verify-before-store contract holds for open-ended reasoning: 88/88 consistency-gated acceptances across all four models, machine-checked formal proof, and reasoning-method transfer at 77/80. Memory selection takes 1.4 microseconds; a full reuse completes in 6-23 ms at 36 mWh. Approximate similarity retrieval selects the wrong item 94.3% of the time on a 4,500-item verified store where exact addressing makes zero errors. The store also serves as working context at a scale no shipped engine matches: a 6,000,000-token movable window on a single 46 GB GPU at flat memory, where vLLM stops at 30,399 tokens and SGLang silently truncates past 32,000. On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits every time. A public testbench with free, rate-limited access accompanies this report: https://corbenic-galahad-bench.hf.space
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.