A Verifiable Search Is Not a Learnable Chain-of-Thought
AuthorsHarsh Patel
Resources
This paper argues that some reasoning problems can be solved by search but still cannot be learned as a neat left-to-right chain of thought—what models can distill is often verification and memorization, not the search itself.
Key results
maximum adapter rank used for distillation
cryptarithm backtracking solver coverage
best held-out accuracy across eleven CoT designs and escalations
cryptarithm accuracy after revealing the full cipher key
What the paper found
In this study, Harsh Patel shows that a verifiable search does not reliably distill into a learnable chain-of-thought, using a deterministic-generator benchmark of nine reasoning tasks and a rank-32 LoRA over NVIDIA’s Nemotron-3-Nano-30B-A3B. The easy forward-computable tasks transfer cleanly, but cryptarithm exposes the core failure: a backtracking solver reaches 0.71 accuracy, while the fine-tuned model stays at 0.01–0.07 across eleven chain-of-thought designs, plus RLVR/GRPO and STaR escalations. The paper’s key mechanism is “verdict-as-token,” where the model computes local arithmetic correctly on 97–100% of elimination lines yet emits elimination verdicts that are logically wrong 16–57% of the time, so the learned trace imitates the shape of search without preserving its state or logic. A controlled intervention confirms causality: revealing the full cipher key turns the same cryptarithm instances from 0.03 to 0.571, showing that forward-derivability, not raw solver coverage, is the binding constraint. The result generalizes across four backbones from 3B to 671B and across fine-tuning and in-context prompting, while the practical escape is not to teach search but to memorize a finite candidate catalog and verify it in-trace, as demonstrated by the competition’s 1st-place solution on the NVIDIA Nemotron model reasoning challenge, which reached a Private LB of 0.92.
Original abstract
It is tempting to assume any task solvable by a short program can be taught to a model as its chain-of-thought: write the steps out, fine-tune, and the model follows. This paper shows the assumption fails for an identifiable class of procedures. The testbed is nine reasoning tasks, each from a deterministic generator; public and hidden splits share generators, so held-out data proxies test accuracy. I reverse-engineer the generators into Python solvers, render them as chain-of-thought, and distill into a rank-<= 32 LoRA over a 30B (3.5B-active) Nemotron model. Forward-computable tasks install readily: lookup/arithmetic and an 8-bit boolean task transfer (>= 0.99 and 0.68). Cryptarithm does not: distilling its backtracking search holds at 0.01-0.07 across eleven chain-of-thought designs, RL from verifiable rewards, and self-training, even though a search solver answers 71% of instances. This is not a capability gap. The model does the arithmetic on 97-100% of lines and ranks the correct cipher in its top eight on 71%; it cannot carry the search forward as a left-to-right derivation. Fine-tuning learns the shape of a verifiable elimination step while its verdicts become unconditional templates, correct only 16-57% of the time ("verdict-as-token"). The ceiling holds across backbones from 3B to 671B and across fine-tuning and prompting; a controlled intervention isolates the cause: revealing the cipher key, which turns the derivation forward, lifts the same instances from 0.03 to 0.57. When a procedure's only solution is search over information-free structure, no faithful forward chain-of-thought exists to imitate. The task becomes learnable only by removing the search, precomputing its combinatorial core into a catalog and reducing the trace to recall plus verification; the 1st-place solution reaches Private LB 0.92 this way. What distills is memorization and verification, not search.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.