NTH

Streaming Communication in Multi-Agent Reasoning

AuthorsZhen Yang, Xiaogang Xu, Wen Wang, Cong Chen, Xander Xu, Ying-Cong Chen

July 8, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that letting multi-agent reasoning systems pass along partial thoughts as they are generated can make them both faster and, surprisingly, more accurate.

Key results

7.3%
Claude avg accuracy gain over Serial

Average improvement across three topologies on Claude Opus 4.6

1.5%
GPT avg accuracy gain over Serial

Average improvement across three topologies on GPT-5.4

22.4%
Peak gain on HMMT 2026

Maximum accuracy gain reported for STREAM MA over Serial

68.2%
A64 accuracy with auto steps

Step-level scaling baseline at A=64 with LLM-decided step count

73.5%
A64 accuracy with S=64

Accuracy after increasing per-agent steps to 64

26.9x
Speedup at A=S=64

Measured wall-clock speedup from step-level scaling

What the paper found

Streaming Communication in Multi-Agent Reasoning introduces STREAM MA, a multi-agent protocol from HKUST(GZ), Alibaba Group, and Zhejiang University that replaces the usual generate-then-transfer rule with step-level forwarding: each reasoning step is streamed downstream immediately, allowing adjacent agents to overlap execution instead of waiting for a full response. The paper gives a closed-form analysis of three execution modes—Single, Serial, and Stream—showing that Stream is optimal when early steps are more reliable than later ones, while also deriving a speedup upper bound and an exact cost ratio under LLM serving assumptions. Empirically, across eight benchmarks spanning AIME 2025, AIME 2026, HMMT 2026, GPQA-Diamond, HLE, and LiveCodeBench, using Claude Opus 4.6 and GPT-5.4 over Chain, Tree, and Graph topologies, STREAM MA improves average accuracy by +7.3 pp on Claude and +1.5 pp on GPT versus Serial, with a peak gain of +22.4 pp on HMMT 2026. The study also reports a step-level scaling law: at A=64, increasing per-agent steps S from auto-decided length to 64 raises accuracy from 68.2% to 73.5% and yields 26.9× wall-clock speedup, reaching 83% of the theoretical 32.3× bound. With A=S=4 on Claude Opus 4.6, the protocol achieves about 2.30× speedup, and under full prefix caching it is 7.5% cheaper than Serial, while the no-cache regime reverses that advantage.

Original abstract

Multi-agent reasoning systems adopt a "generate-then-transfer" paradigm that forces end-to-end latency to scale linearly with pipeline depth. We introduce StreamMA, a multi-agent reasoning system that streams each reasoning step to downstream agents as soon as it is generated, pipelining adjacent agents and thus reducing latency. Surprisingly, this pipelining also improves effectiveness: because multi-step reasoning quality is non-uniform and early steps are more reliable than later ones, working with these reliable early steps instead of the full chain prevents error-prone late steps from misleading downstream agents. We formalize both advantages with the first closed-form joint analysis of stream, serial, and single protocols, deriving the effectiveness ordering, speedup upper bound, and cost ratio. Across eight reasoning benchmarks spanning mathematics, science, and code, two frontier LLMs (Claude Opus 4.6 and GPT-5.4), and three topologies (Chain, Tree, Graph), StreamMA outperforms both baselines (avg. +7.3 pp, max +22.4 pp on HMMT 2026; Claude Opus 4.6-high). Beyond these contributions, we discover a "step-level scaling law": increasing per-agent steps consistently improves both effectiveness and efficiency, a new scaling dimension orthogonal to and composable with agent-count scaling.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis