In-Context World Modeling for Robotic Control
AuthorsSiyin Wang, Junhao Shi, Senyu Fei, Zhaoyang Fu, Li Ji, Jingjing Gong, Xipeng Qiu
Resources
This paper teaches robots to infer how their body and environment work from a short self-generated experience window, so they can adapt to new camera angles or robot setups without retraining.
Key results
Multi-View BC average success on unseen LIBERO viewpoints
ICWM average success on unseen LIBERO viewpoints
Number of in-domain azimuth angles used for training
Number of held-out azimuth angles used for evaluation
Standard multi-view VLA success before viewpoint shift on UR5e
Standard multi-view VLA success after viewpoint shift on UR5e
What the paper found
In Context World Modeling for Robotic Control, from Fudan University and the Shanghai Innovation Institute, reframes generalization failure in Vision-Language-Action systems as test-time system identification: instead of assuming a fixed camera pose or robot morphology, the policy first collects a short self-generated interaction prefix and uses that context to infer the latent configuration before acting. The method, ICWM, reuses a Qwen2.5-VL-3B backbone with FAST action tokenization and requires no parameter updates, task-specific demonstrations, or reward signal at deployment. On LIBERO with 8 training viewpoints and 6 unseen OOD viewpoints, ICWM raises average unseen-view success from 19.8% with Multi-View BC to 25.0%, and the paper reports especially large gains on LIBERO-Long, where OOD success reaches 25.0% versus 19.8%. In real-robot experiments on a 6-DoF UR5e with 12 cameras, standard multi-view performance drops from 68% to 17% under viewpoint shift, while ICWM substantially narrows that gap. Ablations show the interaction prefix matters structurally: removing images causes the largest collapse, and false context performs worse than no context, indicating the model genuinely conditions on inferred world dynamics. The approach also transfers beyond camera shifts, improving under distractor objects, novel table textures, and morphology changes such as rigid spacers of 20, 40, and 80 mm, with a particularly notable recovery at 80 mm where the baseline largely fails.
Original abstract
Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because they are typically conditioned only on current observations and language instructions. By ignoring the underlying system configuration as a variable, these models implicitly assume a fixed execution context encountered during training, necessitating data-intensive fine-tuning for any new environment. In this work, we introduce In-Context World Modeling (ICWM), a framework that treats system identification as an in-context adaptation problem. ICWM enables robot policies to autonomously infer essential system variables from a short history of self-generated, task-agnostic interactions. Unlike traditional In-Context Learning that uses demonstrations to specify what task to perform, ICWM leverages the context window to understand how the system operates. By processing these interactions before task execution, the model implicitly captures the world dynamics of the current system, enabling adaptation to novel configurations without parameter updates. Extensive experiments in simulation and on real-world robot platforms demonstrate that ICWM significantly outperforms standard VLA baselines on novel camera viewpoints.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.