Thinking with Looped Flows
AuthorsAyhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim
Resources
Looped Flows lets models spend more computation at inference time by repeatedly refining noisy predictions, achieving strong results on several challenging reasoning benchmarks.
Key results
Looped flows’ result, compared with TRM’s 44.6% using the same architecture.
Looped flows’ result, compared with TRM’s 7.8%.
Accuracy after increasing inference from 8 steps, where accuracy was 74.5%, to 128 steps.
Share of TRM failure cases resolved by looped flows across roughly 65,000 Sudoku instances.
What the paper found
Thinking with Looped Flows introduces a recurrent reasoning method that combines looped neural networks with continuous flow matching for categorical data. Unlike conventional looped models such as TRM, which use truncated backpropagation and can fail to train early updates for later computation, the method trains a stateful denoiser on locally supervised objectives at progressively lower noise levels, using the same noise-target pair across steps to temporally align recurrent states. At inference, it integrates a probability-flow ODE or stochastic differential equation while updating the hidden state, so additional computation can be purchased through a finer temporal grid, and different initial noise samples can produce multiple valid solutions. This complements autoregressive reasoning approaches associated with OpenAI, which externalize computation as text, by accumulating computation in distributed hidden states. On six reasoning benchmarks, looped flows outperform prior looped models on five: with the same architecture as TRM, ARC-AGI-1 pass@2 rises from 44.6% to 58.8%, while ARC-AGI-2 rises from 7.8% to 12.2%. On Sudoku-Extreme, increasing inference steps from 8 to 128 raises accuracy from 74.5% to 97.9%. In an analysis of roughly 65,000 Sudoku test instances, the method resolves 90.9% of TRM’s failure cases, including non-convergence and spurious attractors. Stochastic integration and best-Q selection using five trajectories further improve accuracy and solution coverage, especially on N-Queens and Graph Coloring, where multiple valid outputs exist.
Original abstract
Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.