NTH

Self-Organizing Agent Teams Learn to Reason Together

AuthorsAneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

AffiliationsStanford University · Together AI · Goizueta Business School, Emory University

October 3, 2026 3 min read
Watch on YouTube
The one-line take

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Key results

66.7%
Math-physics SAT accuracy

Average accuracy across five mathematics and physics benchmarks.

48.8%
Strongest-member accuracy

Average accuracy of the strongest individual team member on the same suite.

59.0%
Routing-oracle coverage

Perfect selection among members’ independent answers.

71.2%
AIME 2026 SAT accuracy

SAT accuracy on AIME 2026.

13.4%
AIME 2026 oracle improvement

SAT’s absolute advantage over routing-oracle coverage on AIME 2026.

0.90
Demonstrability correlation

Spearman correlation between demonstrability and improvement over the strongest member across eight benchmarks.

What the paper found

Self-Organizing Agent Teams, or SAT, treats team organization as a learnable capability rather than a fixed debate or routing protocol. The paper is motivated in part by an OpenAI-reported incident in which agents used shared infrastructure to coordinate, including activity involving Hugging Face, but its experiments focus on controlled reasoning tasks. SAT uses offline evolutionary search to learn reusable conversational strategies governing roles, participation, information flow, challenge, repair, and synthesis; the strategies are then frozen and transferred without problem-specific decomposition. A mathematics-and-physics team combining o3-mini, Claude Sonnet 4, and DeepSeek-V3, trained on 15 AIME 2024 problems, averages 66.7% accuracy across five benchmarks, compared with 48.8% for its strongest member and 59.0% for a perfect routing oracle over independent answers. On AIME 2026, it reaches 71.2% and exceeds that oracle by 13.4 percentage points, showing that interaction can construct solutions absent from individual outputs. A separate team using Gemini-2.5-Flash, Llama-4-Maverick, and GPT-4.1 reaches 87.9% certificate coverage but only 72.8% final accuracy, exposing a selection bottleneck: generating correct reasoning is easier than recognizing it. Across eight benchmarks, this gap is strongly associated with demonstrability, measured by whether external judges can distinguish correct from incorrect reasoning, with Spearman ρ = 0.90. The central result is that learned organizational structure can unlock collaborative computation beyond aggregation, voting, or single-agent scaling.

Original abstract

Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $ρ=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Atria Dawn: The Dawn of Agentic Superintelligence

Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv, Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang, Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li, Jiaqiang Li, Peng Li, Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu, Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan...

Atria Dawn explores how research-oriented AI agents can move beyond completing tasks to partnering with humans on scientific discovery while preserving human oversight.

Read analysis