NTH

Loop the Loopies!

AuthorsZitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai

July 22, 2026 2 min read
Watch on YouTube
The one-line take

Loopie argues that repeatedly looping a sparse MoE Transformer can beat simply scaling model size, reaching remarkable mathematical and scientific reasoning performance.

Key results

20B
Loopie-20B-A2B total parameters

Total parameter count of the larger Loopie model.

2B
Loopie-20B-A2B active parameters

Parameters activated per token in the larger MoE model.

261.53
Layer-loop peak throughput

Peak throughput in TFLOPS/s reported for Loopie-20B-A2B.

189.65
Vanilla baseline peak throughput

Peak throughput in TFLOPS/s for the Qwen3-like 30B-A3B baseline.

92.09
AIME 2024 score

Loopie-20B-A2B Thinking accuracy on AIME 2024.

94.21
AMC score

Loopie-20B-A2B Thinking accuracy on AMC.

What the paper found

“Loop the Loopies!” from IQuest Research introduces Loopie, a pair of Qwen3-MoE-style language models that make recurrent computation competitive with conventional Transformer scaling under matched pre-training time. Loopie-20B-A2B contains 20B total parameters with 2B active per token, while Loopie-6B-A0.6B contains 6B with 0.6B active. Its central innovation is layer-loop recurrence: each attention/MoE layer is applied twice consecutively before the model advances, rather than repeatedly traversing the entire stack as in model-loop systems such as Ouro and Huginn. The hardware-aware Loopie Recipe halves stored depth, doubles the per-device microbatch, and reinvests the measured efficiency gain into width and capacity; the final model reaches 261.53 TFLOPS/s peak throughput versus 189.65 TFLOPS/s for a Qwen3-like 30B-A3B baseline. In matched training, Loopie overtakes that vanilla baseline after roughly 600B tokens. Pre-training uses NVIDIA’s Nemotron-CC-v2-HQ and related Nemotron datasets, followed by 2T tokens of supervised pre-training at large language-model batch scales, then reinforcement learning using GSPO with DAPO-style asymmetric clipping and dynamic filtering. The resulting Loopie-20B-A2B Thinking scores 92.09 on AIME 2024 and 94.21 on AMC, after only 3.5T pre-training tokens compared with 25T for the cited Nemotron models. The paper’s main conclusion is that carefully scheduled recurrence, rather than looping alone, can become a scalable compute-allocation strategy for sparse large language models.

Original abstract

We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N-fold increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. Our novel post-training pipeline equips Loopie with strong reasoning abilities. At the 2025 IMO and IPhO, Loopie achieves gold-medal performance without tools.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis