NTH

Test-Time Scaling via Error Localization

AuthorsRajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta, Deepanway Ghosal, Aravindan Raghuveer

July 29, 2026 2 min read
Watch on YouTube
The one-line take

TTEL makes language-model reasoning more efficient by finding where a solution went wrong and branching from the last correct step instead of starting over.

Key results

71.0%
LiveCodeBench TTEL pass@64

Qwen3-8B performance with TTEL

360.4k
TTEL generated tokens

Average LiveCodeBench tokens at pass@64

735.0k
Independent sampling tokens

Average LiveCodeBench tokens at pass@64

0.820
AIME-25 TTEL pass@16

Best reported TTEL scaling result on AIME-25

486.0
Unfiltered first-turn spikes

Spike count after removing null-baseline filtering, compared with 19.3

What the paper found

Researchers at Google DeepMind introduce Test-Time Scaling via Error Localization, or TTEL, an inference-only search algorithm for reasoning models such as Qwen3-8B and Qwen3-4B-Thinking-2507. Instead of discarding failed solutions or regenerating them from scratch, TTEL re-scores the original token sequence with task feedback and with a non-diagnostic null prompt, subtracts the resulting probability shifts, and identifies the token with the strongest feedback-specific disagreement. It then truncates the trajectory at that point and branches a new continuation while retaining the valid prefix, creating a token-level prefix tree without gradient updates or a separate reward model. On LiveCodeBench, Qwen3-8B reaches 71.0% pass@64 with 360.4k generated tokens, versus 64.6% and 735.0k tokens for independent sampling. On AIME-25, TTEL reaches 0.820 pass@16 and generalizes to HMMT-25 and the smaller Qwen3-4B-Thinking-2507. Ablations show that preserving the full reasoning trace and using execution feedback improve localization, while removing null-baseline filtering increases detected first-turn spikes from 19.3 to 486.0 and reduces performance to 0.592 pass@k. The central result is a compute-efficiency frontier in which inference is concentrated on correcting the suspected faulty suffix rather than repeating valid reasoning.

Original abstract

Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn refinement operate without token-level credit assignment, resulting in computational inefficiency, since valid reasoning prefixes are frequently discarded. In this work, we introduce Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localization. By comparing conditional probabilities under informed feedback against a null-context baseline, TTEL isolates the step at which an error occurred. The algorithm then truncates the trajectory and branches a new generation, maximally reusing the valid prefix. Extensive evaluations demonstrate that TTEL establishes strictly dominating Pareto frontiers across sequential reasoning domains, measured by pass-at-k vs. generated-token cost. With Qwen3-8B on LiveCodeBench, TTEL attains a pass@64 of 71.0% while generating approximately half as many tokens as independent sampling (360.4k vs. 735.0k). Generalizing to math benchmarks AIME-2025 and HMMT-2025, TTEL cleanly outperforms competing test-time baselines across both Qwen3-8B and Qwen3-4B-Thinking-2507.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis