NTH

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

AuthorsSietse Schelpe

August 5, 2026 3 min read
Watch on YouTube
The one-line take

Instead of making the model smarter, this system remembers verified answers so it can return them instantly, exactly, and forever.

Key results

100%
Verified reuse accuracy

Accuracy across 180 fresh instances spanning nine problem families

0
Generation tokens per reuse

Verified answers are produced without token generation

23
Maximum reuse latency

Full reuse completes within the reported 6–23 ms range

94.3%
Approximate retrieval error

Wrong-item selection rate on a 4,500-item verified store

6M
Movable context window

Token capacity measured on a single 46 GB GPU

What the paper found

An industry report from Sietse Schelpe of Corbenic AI presents Galahad, a proprietary system whose Merlin integrity layer stores solutions only after answer-key-free verification, while keeping the underlying language model frozen. Instead of regenerating answers, Galahad stores verified parameterized methods and executes them deterministically on new instances. Across 180 fresh instances from nine algorithmic and mathematical families, frozen Gemma-4-12B achieved 100% accuracy at 0 generation tokens, and the same result held for Qwen3-14B, DeepSeek-Coder-V2-Lite, and Phi-4, spanning dense and mixture-of-experts architectures. Reuse completed in a measured 6–23 ms, while consistency-gated reasoning achieved 88/88 accepted reuses and transferred methods correctly in 77/80 cases. The report also argues that exact addressing is essential: on a 4,500-item verified store, approximate similarity retrieval selected the wrong item 94.3% of the time, whereas exact addressing produced zero collisions. Galahad further maintained a movable 6,000,000-token context on one 46 GB GPU, compared with vLLM’s 30,399-token ceiling and SGLang’s silent truncation beyond roughly 32,000 tokens. The authors carefully scope the claim: Anthropic’s frontier models and Google DeepMind’s Gemma reporting still show superior cold-start reasoning, but on previously solved and verified workloads, this frozen-model approach trades repeated generation for bit-exact, auditable execution at near-zero marginal cost.

Original abstract

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning nine problem families, four architectures from four vendors - dense and mixture-of-experts - each score 180/180 at zero generation tokens per answer: execution-bound capability decoupled from parameter scaling. A negative control attributes the capability fully to the memory: emptied, it solves nothing. The same verify-before-store contract holds for open-ended reasoning: 88/88 consistency-gated acceptances across all four models, machine-checked formal proof, and reasoning-method transfer at 77/80. Memory selection takes 1.4 microseconds; a full reuse completes in 6-23 ms at 36 mWh. Approximate similarity retrieval selects the wrong item 94.3% of the time on a 4,500-item verified store where exact addressing makes zero errors. The store also serves as working context at a scale no shipped engine matches: a 6,000,000-token movable window on a single 46 GB GPU at flat memory, where vLLM stops at 30,399 tokens and SGLang silently truncates past 32,000. On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits every time. A public testbench with free, rate-limited access accompanies this report: https://corbenic-galahad-bench.hf.space

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis