T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
AuthorsJunyao Yang, Yucheng Shi, Zhongzhi Li, Ruhan Wang, Zongxia Li, Haitao Mi, Leowei Liang
Resources
T1 trains a large MoE agent to complete hundreds of shell interactions, substantially improving long-horizon coding and terminal-task performance through verifier rewards and stable RL techniques.
Key results
Total parameters in the Mixture-of-Experts model.
Synthesized terminal tasks used for dense-reward reinforcement learning.
Reduced from 0.021 using TITO and R3.
T1 score under the shared agent harness.
T1 average reward on the longer-horizon evaluation suite.
What the paper found
This Tencent-backed paper introduces T1, a 122B-total-parameter Mixture-of-Experts agent obtained by reinforcement-learning Qwen3.5-122B-A10B in real cloud sandboxes, where tasks are judged by executable verifiers rather than preference models and can require more than 300 tool-call turns. Its training recipe combines dense process rewards based on absolute passing assertions with PPO, token-in-token-out, or TITO, which preserves the exact sampled token identifiers across multi-turn trajectories, and rollout routing replay, or R3, which reuses the MoE experts selected during inference. Together, TITO and R3 reduce the training-to-inference log-probability gap from 0.021 to 0.013 and produce zero token drift in the loss region. Training uses an out-of-distribution T1-15k corpus of 15,000 synthesized terminal tasks, separate from evaluation data, to test capability transfer rather than benchmark memorization. On Terminal-Bench 2.1, T1 reaches 64.0 percent versus 43.8 percent for the base Qwen3.5 model, surpassing GPT-5.4 at 54.8 percent, DeepSeek-V4-Flash at 56.9 percent, and Claude Opus 4.6 at 63.8 percent under the same harness. It also achieves 27.9 percent average reward on Long-Horizon Terminal-Bench, matching Gemini-3.1-Pro and exceeding GPT-5.4, although the study reports longer interaction costs and remaining failures on the hardest tasks.
Original abstract
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording the sampler's per-token expert choices at every MoE layer and replaying them during training. Third, fully out-of-distribution training corpus: isolated seeds and synthesized tasks disjoint from Terminal-Bench 2.1 ensures gains reflect genuine capability transfer over benchmark overfitting. Together, TITO and R3 cut the training-to-inference log-probability difference from 0.021 to 0.013, with exactly aligned zero token drift in the loss region. On Terminal-Bench 2.1, our post-train pipeline raises initial base model from 43.8% to T1 with 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.