NTH

Thinking with Looped Flows

AuthorsAyhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim

September 15, 2026 2 min read
Watch on YouTube
The one-line take

Looped Flows lets models spend more computation at inference time by repeatedly refining noisy predictions, achieving strong results on several challenging reasoning benchmarks.

Key results

58.8%
ARC-AGI-1 pass@2

Looped flows’ result, compared with TRM’s 44.6% using the same architecture.

12.2%
ARC-AGI-2 pass@2

Looped flows’ result, compared with TRM’s 7.8%.

97.9%
Sudoku-Extreme scaling

Accuracy after increasing inference from 8 steps, where accuracy was 74.5%, to 128 steps.

90.9%
TRM failure recovery

Share of TRM failure cases resolved by looped flows across roughly 65,000 Sudoku instances.

What the paper found

Thinking with Looped Flows introduces a recurrent reasoning method that combines looped neural networks with continuous flow matching for categorical data. Unlike conventional looped models such as TRM, which use truncated backpropagation and can fail to train early updates for later computation, the method trains a stateful denoiser on locally supervised objectives at progressively lower noise levels, using the same noise-target pair across steps to temporally align recurrent states. At inference, it integrates a probability-flow ODE or stochastic differential equation while updating the hidden state, so additional computation can be purchased through a finer temporal grid, and different initial noise samples can produce multiple valid solutions. This complements autoregressive reasoning approaches associated with OpenAI, which externalize computation as text, by accumulating computation in distributed hidden states. On six reasoning benchmarks, looped flows outperform prior looped models on five: with the same architecture as TRM, ARC-AGI-1 pass@2 rises from 44.6% to 58.8%, while ARC-AGI-2 rises from 7.8% to 12.2%. On Sudoku-Extreme, increasing inference steps from 8 to 128 raises accuracy from 74.5% to 97.9%. In an analysis of roughly 65,000 Sudoku test instances, the method resolves 90.9% of TRM’s failure cases, including non-convergence and spurious attractors. Stochastic integration and best-Q selection using five trajectories further improve accuracy and solution coverage, especially on N-Queens and Graph Coloring, where multiple valid outputs exist.

Original abstract

Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis