Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator
AuthorsZihan Wang, Seungjun Lee, Yinghao Xu, Gim Hee Lee
Resources
Image2Sim turns ordinary RGB-D image sequences into large-scale interactive navigation simulators, letting embodied AI agents train in realistic neural worlds instead of expensive hand-built environments.
Key results
constructed from real and synthetic captures
synthetic vision-language-action training data
approximately 40 FPS on a single RTX 4090
Image2Nav with 180° FOV zero-shot in Habitat
Image2Nav with 180° FOV zero-shot in Habitat
Image2Nav with 180° FOV zero-shot in Habitat
What the paper found
Image2Sim, from National University of Singapore and HKUST, reframes embodied navigation as a scaling problem in neural simulation rather than manual annotation. The paper’s core idea is to decouple persistent 3D grounding from photorealistic rendering: a feed-forward feature-Gaussian encoder lifts posed RGB-D captures into explicit geometry and semantics in one pass, then a Geometry-Aware One-Step Pixel Flow renderer completes missing regions with a conditional MeanFlow model guided by alpha maps and DINOv3 features. This produces a real-time simulator that renders panoramic RGB-D at about 40 FPS on a single RTX 4090 while also serving as an automatic embodied data engine. Using RealSee3D, Structured3D, ARKitScenes, HM3D, ScanNet, Gibson, and Matterport3D, it constructs 19,936 interactive scenes and synthesizes over 10M vision-language-action samples. Training Image2Nav only in Image2Sim yields new state-of-the-art zero-shot transfer in Habitat, including 70.3 SR and 65.6 SPL on R2R-CE with 180° FOV, 70.7 SR and 59.1 SPL on RxR-CE, and 53.7 SR and 42.7 SPL on REVERIE-CE. The scaling study shows consistent gains from 46.1 to 66.3 SR as synthetic training grows from the human-annotated baseline to 10M samples, and real-world tests on a Hello Robot Stretch 3 improve path-following success from 8/20 to 11/20 and goal-oriented success from 5/20 to 9/20, indicating that physically executable neural simulation can narrow the sim-to-real gap for embodied navigation.
Original abstract
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.