ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI
AuthorsZhengran Ji, Jonathan Hyun, Boyuan Chen
AffiliationsDepartment of Computer Science, Duke University, Durham, USA · Department of Electrical and Computer Engineering, Duke University, Durham, USA
Resources
ORCH shows that organizing embodied AI agents into task-specific hierarchical teams can substantially improve coordination and performance in complex wildfire-response missions.
Key results
Benchmark missions spanning reconnaissance, rescue, transport, containment, and suppression.
Largest evaluated teams combined firefighters, bulldozers, drones, and helicopters.
Average improvement relative to four representative embodied multi-agent baselines.
Average execution-efficiency improvement measured with performance-curve AUC.
Average improvement relative to the four baseline frameworks.
Average execution-efficiency improvement relative to the baselines.
What the paper found
ORCH, or Organizing Roles and Coordination Hierarchies, treats organizational design as a computational component of embodied collective intelligence. Instead of imposing one fixed topology, it builds task-specific hierarchies from the mission, available workers, and capabilities. Horizontal managers implement pooled interdependence by distributing independent subtasks in parallel, while vertical managers implement sequential interdependence through prerequisite-aware mission phases; these structures can be recursively combined into teams of teams. In 25 CREW-Wildfire missions, teams of up to 50 heterogeneous agents—including firefighters, bulldozers, drones, and helicopters—were evaluated with 8 language models, including OpenAI’s ChatGPT-5.4, Llama-4-Scout, Gemma-4-it, Qwen-3.6, and DeepSeek-V4-Pro, against CAMON, COELA, HMAS-2, and Embodied. Human-designed ORCH hierarchies improved final score by 63.97% and execution efficiency, measured by performance-curve AUC, by 74.29% relative to the four baselines; language-model-generated hierarchies still improved these metrics by 43.63% and 52.53%. A critic-guided generation process reduced unnecessary managerial layers and produced more balanced structures than uncriticized generation. ORCH’s advantage persisted across missions and models, while aggregate collective performance did not increase monotonically with model scale: mid-sized Gemma-4-it and Qwen-3.6 outperformed larger systems such as ChatGPT-5.4 and DeepSeek-V4-Pro in this embodied setting. The central result is that role assignment, authority, information flow, and temporal coordination can matter as much as the capability of individual agents.
Original abstract
Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the physical tasks they perform impose fundamentally different coordination requirements. Here we show that principles from human organization theory can be operationalized to organize large, heterogeneous collectives of embodied artificial agents. We introduce ORCH (Organizing Roles and Coordination Hierarchies), which constructs task-specific hierarchical organizations by combining pooled interdependence for work that can proceed concurrently with sequential interdependence for work governed by prerequisite relationships. Across 25 wildfire-response missions spanning reconnaissance, rescue, transportation, resource management, containment and suppression, we evaluated teams of up to 50 heterogeneous agents using eight large language models. Organizations constructed using these principles consistently outperformed four representative embodied multi-agent approaches across mission outcome, execution efficiency, exploration and computational resource use. Human-designed ORCH organizations improved final score by 63.97% and execution efficiency by 74.29% on average relative to the four prior frameworks. Organizations generated automatically by language models improved these measures by 43.63% and 52.53%, respectively. These advantages persisted across missions and underlying language models. Notably, collective performance was not monotonically determined by model scale. Analysis of long-horizon missions showed that hierarchical organization enabled teams to preserve concurrent activity within specialized groups while coordinating ordered transitions between mission phases.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.