NTH

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

AuthorsNadun Ranawaka, Josiah Wong, Wei-Lin Pai, Wei-Teng Chu, Tianyuan Dai, Masoud Moghani, Hang Yin, Yunfan Jiang, Wesley Durbano, Brandon Huynh, Yu Fang, Linxi Fan, Danfei Xu, Ruohan Zhang, Li Fei-Fei, Bowen Wen, Ajay Mandlekar, Yuke Zhu

June 29, 2026 2 min read
Watch on YouTube
The one-line take

SimFoundry turns a single real video into editable simulation twins that help robot policies train, generalize, and transfer to the real world more reliably.

Key results

12
reconstruction scenes

Number of table-top scenes used to evaluate zero-shot reconstruction fidelity

0.81-0.92
zero-shot F1

SimFoundry zero-shot geometric reconstruction F1 range across scenes

0.93-0.99
tuned F1

F1 range after 3 minutes of per-object tuning

0.911
Pearson correlation

Mean real-to-sim correlation across 7 tasks and 5 policy architectures

0.018
MMRV

Mean maximum rank violation for SimFoundry real-to-sim evaluation

40%
task success improvement

Average real-world success gain from task cousins

What the paper found

SimFoundry, from NVIDIA with collaborators at Stanford, Georgia Tech, and UT Austin, is a modular real-to-sim pipeline that turns a single RGB video into an interactive, sim-ready digital twin and then expands it into affordance-preserving object, scene, and task cousins for robot policy learning and evaluation. The system decomposes reconstruction into extraction, mesh generation, pose alignment, articulation inference, physics annotation, and PyBullet stabilization, while using foundation models such as DepthAnything3, SAM3, FoundationPose, Gemini-Pro-3, Hunyuan3D 2.1, and TRELLIS.2. Across 12 reconstructed tabletop scenes, it reaches zero-shot F1 scores of 0.81–0.92 and improves to 0.93–0.99 with only 3 minutes of per-object tuning; compared with SAM3D, it lowers reconstruction error and raises geometric fidelity. For real-to-sim evaluation across 7 manipulation tasks and 5 policy architectures, SimFoundry’s simulated rankings track real-world success with mean Pearson correlation 0.911 and mean MMRV 0.018, exceeding the PolaRiS baseline by more than 0.59 Pearson points. For sim-to-real training, object cousins, scene cousins, and task cousins improve average real-world task success by 17%, 21%, and 40%, respectively, while multi-task cousin data raises generalist policy performance by up to 31% in simulation and 18% in the real world, including 29% success on held-out tasks. The automatic background pipeline also outperforms manual alignment, with PSNR 15.29 versus 12.91 and NCC 0.749 versus 0.549.

Original abstract

Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis