When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
AuthorsKaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung
AffiliationsUC Berkeley · University of Washington · Princeton University · Bespoke Labs
Resources
This paper shows that LLM agents often get worse returns from prolonged thinking and can perform better by splitting compute across multiple shorter attempts.
Key results
Tasks drawn from FrontierCS, ALE-Bench, FlashInfer-Bench, and MLS-Bench.
Elo gained per decade of compute.
Token budget used for the headline agent evaluation.
FrontierCS Polyomino Packing budget where marginal scaling matches sampling.
Elo improvement over one 100M-token session.
Elo improvement over ten short sessions at the same 100M-token total budget.
What the paper found
This paper introduces Elo-per-token, a compute-aware metric for open-ended tasks that tracks each session’s best-so-far solution at successive token budgets, compares submissions within each task, and fits a Bradley–Terry model to produce Elo curves comparable across heterogeneous score scales. The study evaluates 14 tasks from FrontierCS, ALE-Bench, FlashInfer-Bench, and MLS-Bench using Kimi K2.7, OpenAI’s GPT-5.5 in Codex, Anthropic’s Claude Opus 4.8 in Claude Code, and Google’s Gemini 3.5 Flash in Gemini CLI, with sessions reaching 100M tokens. Independent sampling provides a distribution-free baseline of 400 Elo per decade of compute: agents often outperform this slope early, but their marginal gains decline and eventually fall below it, indicating that long trajectories lose efficiency after context compactions. Historical AtCoder Heuristic Contest contestants instead improve superlinearly over days, suggesting continual learning and persistent headroom beyond current agents. Specialized strategies, including AdaEvolve, GEPA, and TTT-Discover with gpt-oss-20b, also show diminishing returns rather than human-like scaling. The authors define a scaling inflection point where an agent’s marginal Elo gain matches the 400-Elo reference; on FrontierCS Polyomino Packing, this point occurs at 38M tokens. Splitting a 100M-token budget into three sessions, rather than one long session or ten short sessions, gains 264 Elo over one session and 355 Elo over ten, demonstrating a practical rule for compute allocation.
Original abstract
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. We apply it to four general-purpose agents on four open-ended benchmarks, with sessions of up to 100M tokens, and to three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference, for which Elo grows linearly with log compute. Against this reference, agents can initially convert tokens into Elo faster than independent sampling, but their marginal gains diminish and eventually fall below the reference. In contrast, the strongest historical human contestants improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, providing evidence of continual learning and substantial headroom after agents slow down. We define the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference. Using this point as the per-session budget, we split 100M tokens across parallel sessions on FrontierCS Polyomino Packing, gaining +264 Elo over one long session and +355 over ten short sessions.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.