NTH

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

AuthorsZihan Wang, Seungjun Lee, Yinghao Xu, Gim Hee Lee

July 13, 2026 3 min read
Watch on YouTube
The one-line take

Image2Sim turns ordinary RGB-D image sequences into large-scale interactive navigation simulators, letting embodied AI agents train in realistic neural worlds instead of expensive hand-built environments.

Key results

19936
interactive scenes

constructed from real and synthetic captures

10M
navigation samples

synthetic vision-language-action training data

40
rendering speed

approximately 40 FPS on a single RTX 4090

70.3%
R2R-CE SR

Image2Nav with 180° FOV zero-shot in Habitat

70.7%
RxR-CE SR

Image2Nav with 180° FOV zero-shot in Habitat

53.7%
REVERIE-CE SR

Image2Nav with 180° FOV zero-shot in Habitat

What the paper found

Image2Sim, from National University of Singapore and HKUST, reframes embodied navigation as a scaling problem in neural simulation rather than manual annotation. The paper’s core idea is to decouple persistent 3D grounding from photorealistic rendering: a feed-forward feature-Gaussian encoder lifts posed RGB-D captures into explicit geometry and semantics in one pass, then a Geometry-Aware One-Step Pixel Flow renderer completes missing regions with a conditional MeanFlow model guided by alpha maps and DINOv3 features. This produces a real-time simulator that renders panoramic RGB-D at about 40 FPS on a single RTX 4090 while also serving as an automatic embodied data engine. Using RealSee3D, Structured3D, ARKitScenes, HM3D, ScanNet, Gibson, and Matterport3D, it constructs 19,936 interactive scenes and synthesizes over 10M vision-language-action samples. Training Image2Nav only in Image2Sim yields new state-of-the-art zero-shot transfer in Habitat, including 70.3 SR and 65.6 SPL on R2R-CE with 180° FOV, 70.7 SR and 59.1 SPL on RxR-CE, and 53.7 SR and 42.7 SPL on REVERIE-CE. The scaling study shows consistent gains from 46.1 to 66.3 SR as synthetic training grows from the human-annotated baseline to 10M samples, and real-world tests on a Hello Robot Stretch 3 improve path-following success from 8/20 to 11/20 and goal-oriented success from 5/20 to 9/20, indicating that physically executable neural simulation can narrow the sim-to-real gap for embodied navigation.

Original abstract

Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis