Streaming Communication in Multi-Agent Reasoning
AuthorsZhen Yang, Xiaogang Xu, Wen Wang, Cong Chen, Xander Xu, Ying-Cong Chen
Resources
This paper shows that letting multi-agent reasoning systems pass along partial thoughts as they are generated can make them both faster and, surprisingly, more accurate.
Key results
Average improvement across three topologies on Claude Opus 4.6
Average improvement across three topologies on GPT-5.4
Maximum accuracy gain reported for STREAM MA over Serial
Step-level scaling baseline at A=64 with LLM-decided step count
Accuracy after increasing per-agent steps to 64
Measured wall-clock speedup from step-level scaling
What the paper found
Streaming Communication in Multi-Agent Reasoning introduces STREAM MA, a multi-agent protocol from HKUST(GZ), Alibaba Group, and Zhejiang University that replaces the usual generate-then-transfer rule with step-level forwarding: each reasoning step is streamed downstream immediately, allowing adjacent agents to overlap execution instead of waiting for a full response. The paper gives a closed-form analysis of three execution modes—Single, Serial, and Stream—showing that Stream is optimal when early steps are more reliable than later ones, while also deriving a speedup upper bound and an exact cost ratio under LLM serving assumptions. Empirically, across eight benchmarks spanning AIME 2025, AIME 2026, HMMT 2026, GPQA-Diamond, HLE, and LiveCodeBench, using Claude Opus 4.6 and GPT-5.4 over Chain, Tree, and Graph topologies, STREAM MA improves average accuracy by +7.3 pp on Claude and +1.5 pp on GPT versus Serial, with a peak gain of +22.4 pp on HMMT 2026. The study also reports a step-level scaling law: at A=64, increasing per-agent steps S from auto-decided length to 64 raises accuracy from 68.2% to 73.5% and yields 26.9× wall-clock speedup, reaching 83% of the theoretical 32.3× bound. With A=S=4 on Claude Opus 4.6, the protocol achieves about 2.30× speedup, and under full prefix caching it is 7.5% cheaper than Serial, while the no-cache regime reverses that advantage.
Original abstract
Multi-agent reasoning systems adopt a "generate-then-transfer" paradigm that forces end-to-end latency to scale linearly with pipeline depth. We introduce StreamMA, a multi-agent reasoning system that streams each reasoning step to downstream agents as soon as it is generated, pipelining adjacent agents and thus reducing latency. Surprisingly, this pipelining also improves effectiveness: because multi-step reasoning quality is non-uniform and early steps are more reliable than later ones, working with these reliable early steps instead of the full chain prevents error-prone late steps from misleading downstream agents. We formalize both advantages with the first closed-form joint analysis of stream, serial, and single protocols, deriving the effectiveness ordering, speedup upper bound, and cost ratio. Across eight reasoning benchmarks spanning mathematics, science, and code, two frontier LLMs (Claude Opus 4.6 and GPT-5.4), and three topologies (Chain, Tree, Graph), StreamMA outperforms both baselines (avg. +7.3 pp, max +22.4 pp on HMMT 2026; Claude Opus 4.6-high). Beyond these contributions, we discover a "step-level scaling law": increasing per-agent steps consistently improves both effectiveness and efficiency, a new scaling dimension orthogonal to and composable with agent-count scaling.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.