LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
AuthorsJian Yang, Shawn Guo, Wei Zhang, Tianyu Zheng, Yaxin Du, Haau-Sing Li, Jiajun Wu, Yue Song, Yan Xing, Qingsong Cai, Zelong Huang, Chuan Hao, Ran Tao, Xianglong Liu, Wayne Xin Zhao, Mingjie Tang, Weifeng Lv, Ming Zhou, Bryan Dai
Resources
LoopCoder-v2 shows that, for parallel loop transformers, two loops are the sweet spot: enough extra computation to boost code and agent performance, but not so much that positional mismatch overwhelms the gains.
Key results
mixed text-and-code pretraining corpus
LoopCoder-v2 parameter scale
R=1 baseline score
best 2-loop score
R=1 baseline score
best 2-loop score
What the paper found
LoopCoder-v2, developed by the Beihang University and IQuest Research team, studies how Parallel Loop Transformers saturate under different loop counts for test-time compute scaling. The paper trains a 7B coder from scratch on 18T mixed text-and-code tokens and shows a sharply non-monotonic effect: the 2-loop model is the best operating point, lifting SWE-bench Verified from 43.0% to 64.4% and Multi-SWE from 14.0% to 31.0%, while 3 and 4 loops regress below the baseline. The core technical explanation is a gain–cost trade-off: Cross-Loop Position Offset enables near-single-pass latency and nearly constant KV-cache memory via shared first-loop cache plus gated sliding-window attention, but it injects a fixed positional mismatch at every loop boundary. Internal diagnostics show that loop 2 carries the largest hidden-state update, the largest inter-loop attention shift, and the peak effective rank, whereas later loops become oscillatory, redundant, and increasingly dominated by the offset cost. On broader code and agentic benchmarks, the 2-loop model reaches 84.1 on HumanEval+, 73.9 on MultiPL-E, 46.1 on BigCodeBench-Full, 35.4 on LiveCodeBench, and 40.1 on SWE-bench Verified, demonstrating that a single extra latent refinement loop can outperform far larger models when the loop count is chosen correctly.
Original abstract
Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory with the loop count. Parallel loop Transformers (PLT) alleviate this cost through cross-loop position offsets (CLP) and shared-KV gated sliding-window attention, making loop count a practical design choice. We therefore study PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary. We instantiate this study by training LoopCoder-v2, a family of 7B PLT coders with different loop counts, from scratch on 18T tokens, followed by matched instruction tuning and evaluation. Empirically, the two-loop variant delivers broad gains over the non-looped baseline across code generation, code reasoning, agentic software engineering, and tool-use benchmarks, improving SWE-bench Verified from 43.0 to 64.4 points and Multi-SWE from 14.0 to 31.0 points. In contrast, variants with three or more loops regress, revealing a strongly non-monotonic loop-count effect. Our diagnostics show that loop 2 provides the main productive refinement, while later loops yield diminishing, oscillatory updates and reduced representational diversity. Because the CLP-induced mismatch remains roughly fixed as refinement gains shrink, the offset cost increasingly dominates. This gain--cost trade-off explains PLT's saturation at two loops and provides diagnostics for loop-count selection.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.