DeepLoop: Depth Scaling for Looped Transformers
AuthorsShuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
Resources
DeepLoop makes Transformers deeper by reusing the same blocks repeatedly and adjusts residual scaling to keep this recurrent computation stable.
Key results
DeepLoop increases the conservative residual-scaling exponent from 0.25 to 0.5.
GPT-2 experiments were trained on 50B tokens from FineWeb-Edu.
Validation loss improvement in nats at 7 loops.
Best eight-task lm-evaluation-harness average at 7 loops.
What the paper found
DeepLoop, from researchers at Princeton University and UCLA, addresses a stability problem in looped Transformers, where a small set of physical blocks is revisited to create greater effective depth without storing new parameters. The authors show that shared residual updates are both aggregated across visits and reread across visits, producing a visit-alignment factor that standard DeepNorm analysis misses. In the worst-case aligned regime, the required scaling exponent rises from 0.25 to 0.5, leading to the parameter-free Post-LN rule α=(2N)^0.5 and β=(8N)^−0.5. On OpenAI’s GPT-2 small and GPT-2 medium backbones trained for 50B tokens on FineWeb-Edu, DeepLoop is neutral at one loop and improves validation loss as recurrence increases; at GPT-2 medium with 7 loops, it reduces loss by 0.0278 nats relative to the baseline. The method also improves the eight-task lm-evaluation-harness average, reaching 55.20% in the 1-shot GPT-2 medium setting at 7 loops. Applied to the Hierarchical Reasoning Model on ARC-AGI-1, DeepLoop raises two-vote accuracy from 36.50% to 39.75%. The experiments used NVIDIA H200 GPUs, and a small GPT-2 exponent sweep found reliable training only at 0.5 or above, supporting the theoretical threshold.
Original abstract
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient $κ_R$. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from $1/4$ to $1/2$ as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets $α=(2N)^{1/2}$ and $β=(8N)^{-1/2}$ for unrolled depth $N$. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.