Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
AuthorsZhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, Wenjia Xu
Resources
SIS-Bench tests whether UAV-focused multimodal models understand not just the world around them, but also their own position, motion, and evolving state within it.
Key results
Benchmark size across 13 UAV spatial cognition and self-awareness tasks.
Video sources used to construct SIS-Bench.
Includes proprietary and open-source models such as Gemini 3 Flash, GPT-5.4, Doubao-Seed, Qwen, Kimi, and InternVL.
Human upper-bound performance on SIS-Bench.
Highest overall MLLM accuracy reported on SIS-Bench.
Downstream UAV navigation accuracy, compared with 71.2% for the Qwen2.5-VL 3B baseline.
What the paper found
The paper introduces SIS-Bench, a benchmark that evaluates UAV embodied intelligence through a unified “self-in-space” formulation: models must understand both the external environment and the drone’s own motion. It organizes 4,856 question–answer pairs from 1,646 real-world UAV videos into 13 tasks spanning spatial cognition and self-awareness across perception, memory, and reasoning. Evaluating 26 video-capable multimodal large language models—including Google’s Gemini 3 Flash, OpenAI’s GPT-5.4, ByteDance’s Doubao-Seed models, Qwen, Kimi, InternVL, and GLM—reveals that models understand scene content substantially better than their own action history and ego-motion, with performance degrading as temporal integration and reasoning demands increase. Human accuracy reaches 91.7%, while the best model achieves 71.6%. To address this gap, the authors develop SIS-Motion, which fuses visual embeddings with optical-flow features through a dual-encoder architecture and lightweight connector, trained with SIS-Motion-54K. On SIS-Bench, motion-aware modeling raises spatial accuracy to 74.2% and self-awareness accuracy to 63.7%, improving especially perception and memory, although long-horizon planning remains difficult. The approach also transfers to downstream UAV navigation, reaching 92.2% accuracy compared with 71.2% for the Qwen2.5-VL 3B baseline, demonstrating that explicit ego-motion representations can improve both environmental understanding and flight decision-making.
Original abstract
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification. Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels. Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks. Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.