NTH
AI research

Seeing Together:Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

AuthorsKunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Hao Shi, Yi Zhou, M. Saquib Sarfraz, Danda Pani Paudel, Luc Van Gool

June 30, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches multimodal AI to reason across multiple robots’ first-person views, improving cooperative spatial understanding in both simulation and real-world tests.

Key results

114227
EgoTeam QA pairs

multi-robot egocentric QA dataset size

2326
real-world test QAs

two-quadruped robot real-world evaluation set

70.55%
Habitat accuracy

best SP-CoR average accuracy on Habitat

70.82%
iGibson accuracy

best SP-CoR average accuracy on iGibson

3.87%
improvement over strongest baseline

gain on Habitat over the strongest non-SP-CoR baseline

7.12%
improvement over strongest baseline

gain on iGibson over the strongest non-SP-CoR baseline

What the paper found

Seeing Together introduces CoopSR, the first benchmark for multi-robot cooperative egocentric spatial reasoning, and EgoTeam, a 114,227-question dataset spanning 19 QA types, four difficulty tiers, and teams of 2, 3, or 4 robots in Habitat and iGibson, plus a real-world two-quadruped test set of 2,326 QA pairs. The task pushes multimodal large language models beyond single-view egocentric understanding: they must integrate synchronized robot videos to answer spatial, temporal, visibility, and coordination questions about shared scenes. The proposed SP-CoR framework, built on Qwen2.5-VL and evaluated against 22 MLLM baselines, combines a training-free query-guided spectral frame sampler, spectral- and physics-informed multi-robot fusion, and physics-aligned prompt-space distillation that uses privileged pose supervision only during training. On CoopSR, the strongest variant reaches 70.55% average accuracy on Habitat and 70.82% on iGibson, beating the strongest non-SP-CoR baseline by 3.87 and 7.12 percentage points, respectively. In cross-team-size evaluation with unseen N=4 settings, SP-CoR also generalizes better, and ablations show that removing the fusion module or the distillation path causes substantial drops, confirming that cooperative reasoning depends on explicit geometric alignment rather than simple frame selection or retrieval.

Original abstract

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study this problem through multi-robot cooperative dynamic spatial reasoning, where a model must answer spatial, temporal, visibility, and coordination questions by integrating synchronized egocentric videos from a team of moving robots. To support this setting, we introduce CoopSR, the first benchmark for this task, together with EgoTeam, a multi-robot egocentric QA dataset. EgoTeam contains 114,227 QA pairs spanning 19 question types, four difficulty tiers, and three team sizes in Habitat and iGibson, along with a real-world test set of around 2,326 QAs collected using two quadruped robots. We further propose SP-CoR (Spectral and Physics-Informed Cooperative Reasoner), an MLLM framework for fine-grained cooperative spatial reasoning. SP-CoR combines dynamics-aware multi-robot frame sampling, spectral- and physics-guided view fusion, and physics-aligned prompt distillation, enabling the model to benefit from privileged robot-pose supervision during training while requiring only egocentric videos at test time. Across 22 MLLM baselines, SP-CoR consistently improves cooperative reasoning, outperforming the strongest fine-tuned baseline by +3.87% on Habitat and +7.12% on iGibson. It also shows stronger generalization to unseen team sizes and real-world robot tests. Code can be found at https://github.com/KPeng9510/seeing-together.git.

Read the original paper