Test-Time Scaling via Error Localization
AuthorsRajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta, Deepanway Ghosal, Aravindan Raghuveer
Resources
TTEL makes language-model reasoning more efficient by finding where a solution went wrong and branching from the last correct step instead of starting over.
Key results
Qwen3-8B performance with TTEL
Average LiveCodeBench tokens at pass@64
Average LiveCodeBench tokens at pass@64
Best reported TTEL scaling result on AIME-25
Spike count after removing null-baseline filtering, compared with 19.3
What the paper found
Researchers at Google DeepMind introduce Test-Time Scaling via Error Localization, or TTEL, an inference-only search algorithm for reasoning models such as Qwen3-8B and Qwen3-4B-Thinking-2507. Instead of discarding failed solutions or regenerating them from scratch, TTEL re-scores the original token sequence with task feedback and with a non-diagnostic null prompt, subtracts the resulting probability shifts, and identifies the token with the strongest feedback-specific disagreement. It then truncates the trajectory at that point and branches a new continuation while retaining the valid prefix, creating a token-level prefix tree without gradient updates or a separate reward model. On LiveCodeBench, Qwen3-8B reaches 71.0% pass@64 with 360.4k generated tokens, versus 64.6% and 735.0k tokens for independent sampling. On AIME-25, TTEL reaches 0.820 pass@16 and generalizes to HMMT-25 and the smaller Qwen3-4B-Thinking-2507. Ablations show that preserving the full reasoning trace and using execution feedback improve localization, while removing null-baseline filtering increases detected first-turn spikes from 19.3 to 486.0 and reduces performance to 0.592 pass@k. The central result is a compute-efficiency frontier in which inference is concentrated on correcting the suspected faulty suffix rather than repeating valid reasoning.
Original abstract
Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn refinement operate without token-level credit assignment, resulting in computational inefficiency, since valid reasoning prefixes are frequently discarded. In this work, we introduce Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localization. By comparing conditional probabilities under informed feedback against a null-context baseline, TTEL isolates the step at which an error occurred. The algorithm then truncates the trajectory and branches a new generation, maximally reusing the valid prefix. Extensive evaluations demonstrate that TTEL establishes strictly dominating Pareto frontiers across sequential reasoning domains, measured by pass-at-k vs. generated-token cost. With Qwen3-8B on LiveCodeBench, TTEL attains a pass@64 of 71.0% while generating approximately half as many tokens as independent sampling (360.4k vs. 735.0k). Generalizing to math benchmarks AIME-2025 and HMMT-2025, TTEL cleanly outperforms competing test-time baselines across both Qwen3-8B and Qwen3-4B-Thinking-2507.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.