NTH

Decoding Looped Transformers Better for (Almost) Free

AuthorsWeihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

AffiliationsApple

October 9, 2026 2 min read
Watch on YouTube
The one-line take

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Key results

73.33%
Ouro-2.6B-Thinking AIME 2024 pass@1 with LoopCD

Improves from 61.88% at the unguided baseline.

31.71%
Huginn HumanEval pass@1 with LoopCD-Hidden

Improves from 22.56% at full recurrent depth.

48.2%
Maximum forward FLOPs reduction

Achieved in a half-iteration setting while matching or exceeding the full-depth baseline.

0
LoopCD-Hidden added output-pass overhead

Combines hidden states before the output layers, which are evaluated once.

What the paper found

An Apple-associated study introduces LoopCD, a training-free decoding method that uses a looped Transformer’s early recurrent state as a weaker reference and contrasts it with the final state to guide token selection. LoopCD-Logits applies the contrast to token logits, requiring one extra output pass; LoopCD-Hidden combines hidden states before the output layers, adding 0 output-pass overhead. Across Ouro, Huginn, Parcae, and Looped-Qwen3, the method improves reasoning, code generation, and multiple-choice results without auxiliary models or additional training. On Ouro-2.6B-Thinking, LoopCD raises AIME 2024 pass@1 from 61.88% to 73.33%; on Huginn, LoopCD-Hidden lifts HumanEval pass@1 from 22.56% to 31.71%. An adaptive variant scales guidance according to the probability margin between the top two token choices, focusing updates on uncertain decisions. The key efficiency result is that LoopCD can halve recurrent iterations while matching or exceeding full-depth unguided accuracy, reducing forward FLOPs by 22.5% to 48.2%. The study also evaluates Looped-Qwen3, built from Qwen3-4B, showing that intermediate recurrent states can serve as useful guidance rather than being discarded during decoding.

Original abstract

Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis
03Efficiency

Disaggregated Quantization: Specializing LLM Prefill and Decode

Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh

Disaggregated quantization gives LLM prefill and decode their own specialized weights and formats, improving low-bit accuracy while speeding up first-token generation.

Read analysis