NTH
Research collection

Embodied AI research

Research on intelligence grounded in action and interaction with an environment. Explore navigation, manipulation, and the transfer from simulation to the real world.

48 papers · Latest edition October 4, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Embodied AI papers

Newest editions first.

01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis
04Embodied Ai

Agent as Policy for Robotic Manipulation

Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang

A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.

Read analysis
05Embodied Ai

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Jinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim, Hanjung Kim, Seon Joo Kim

HuRo turns massive human-video collections into robot-training data, substantially improving the robustness and performance of vision-language-action policies.

Read analysis
07Embodied Ai

DriveZero: End-to-End Driving Beyond Human Demonstrations

Hao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang, Kuangye Chen, Xiongjie Zhang, Wentao Dai, Hengchen Dai, Siyuan Liu, Zehao Huang, Naiyan Wang

DriveZero teaches autonomous cars to drive beyond human examples by combining foundation-model perception with reinforcement learning in simulated interactive worlds.

Read analysis
08Embodied Ai

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

Kian Hosseinkhani, Qinhe Peng, George Shramko, Mehran Aghabozorgi, Jianing Qian, Tristan Engst, Alireza Moazeni, Dinesh Jayaraman, Ke Li

IMLE-VLA replaces slow iterative action sampling with a single-step multimodal generator, making vision-language-action robots faster, smoother, and more effective.

Read analysis
09Embodied Ai

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao

OpenWAM turns world-action model training into an open, modular science experiment and shows how video knowledge and robot experience can combine to improve control across simulations and real robots.

Read analysis
10Embodied Ai

LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

Jin Lou, Zhiyuan Jing, Andong Chen, Xupeng Wang, Yuan Xu, Yuexuan Li, Xingdong Zhu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Ying Chu, Jingxuan Zhu, Jingyi Li, Liangliang Chen, Jinyan Liu, Zhiqi Song, Jidong Zhang, Hongming Li, Yuchen Zhu

LM-X helps generalist robots act more reliably by predicting what task stage they are in, what event comes next, and how uncertain their movements are.

Read analysis
16Embodied Ai

G0.5: One Autoregressive Stream for Robot Reasoning and Action

Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao

G0.5 turns a vision-language model into a single-stream robot brain that explains what it is doing while directly generating the actions to do it.

Read analysis
17Embodied Ai

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

Zhenjie Yang, Xingyu Jiao, Guopeng Zhong, Shuzhe Yang, Shi Che, Chao Wu, Chenyu Jiang, Dongjie Zhang, Yideng Zhang, Zheng Zhang, Muyun Jiang, Haisheng Su, Shuang Jin, Donghang Zhang, Chao Yang, Li Chen, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan

HandEdit turns abundant human hand videos into training data for dexterous robots by benchmarking how well image-editing models can adapt them to different robotic hands.

Read analysis
19Embodied Ai

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu

ACE-Data-0 turns everyday homes into multisensory recording studios for teaching embodied AI how humans see, move, touch, and complete long-horizon tasks.

Read analysis
20Embodied Ai

Pictura: Perspective-View Self-Play at Scale for Driving

Yuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut, Victor Bares, Yihong Xu, Éloi Zablocki, Alexandre Boulch, Thibault Buhet, Andrei Bursuc, Matthieu Cord

Pictura trains autonomous driving agents through massive visual self-play, teaching them to handle realistic camera views without relying on privileged state information.

Read analysis
21Embodied Ai

Visual Grounding in Zero-Shot Vision-Language Control

J. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà

The study shows that many vision-language controllers succeed without truly seeing, while selective hazard assistance plus deterministic perception offers a safer path to grounded control.

Read analysis
22Embodied Ai

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Ye Wang, Pei Lin, Xiong-Hui Chen, Haoqi Yuan, Zhixuan Liang, Yiyang Huang, Anzhe Chen, Zixing Lei, Jie Zhang, Tao Zhang, Haoyang Li, Tong Zhang, Chenxi Xiao, Ziyuan Jiao, Qin Jin

Ego2Robot turns everyday human manipulation videos into large-scale synthetic robot experience to help robots generalize beyond their training environments.

Read analysis
23Embodied Ai

HumanCLAW: Can Vision-Language Models Act Through a Body?

Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo

HumanCLAW tests whether vision-language models truly know how to act in the physical world and finds that they struggle mainly to track their own body and progress.

Read analysis
26Embodied Ai

Robostral Navigate

Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal

Robostral Navigate uses an 8B vision-language model and only monocular RGB images to achieve state-of-the-art robot navigation while dramatically reducing training costs.

Read analysis
27Embodied Ai

Robots Acquire Manipulation Skills in Seconds from a Single Human Video

Guangyan Chen, Meiling Wang, Te Cui, Zichen Zhou, Qi Shao, Shalfun Li, Hang Su, Roy Gan, Hao Wang, Mengyin Fu, Yi Yang, Yufeng Yue

HOST lets robots learn new manipulation skills from a single human video in under a minute while preserving what they already know.

Read analysis
28Embodied Ai

UniVR: Thinking in Visual Space for Unified Visual Reasoning

Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin

UniVR teaches AI to reason, understand physics, and plan long tasks directly from visual experience without relying on language supervision.

Read analysis
29Embodied Ai

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou

Xiaomi-Robotics-1 scales robot learning with over 100K hours of language-annotated real-world data to produce a stronger and more adaptable general-purpose manipulation policy.

Read analysis
30Embodied Ai

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li

Xiaomi-Robotics-U0 uses a large multimodal world foundation model to generate consistent robot-centered scenes and videos while improving real-world manipulation performance.

Read analysis
31Embodied Ai

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li

Deform360 is a big real-world dataset that helps robots learn how squishy objects move by combining many camera views, tactile sensing, and dense motion tracking.

Read analysis
32Embodied Ai

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Zihan Wang, Seungjun Lee, Yinghao Xu, Gim Hee Lee

Image2Sim turns ordinary RGB-D image sequences into large-scale interactive navigation simulators, letting embodied AI agents train in realistic neural worlds instead of expensive hand-built environments.

Read analysis
33Embodied Ai

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Haoxiang Ma, Junhao Cai, Xiaoxu Xu, Hao Li, Yuyin Yang, Yang Tian, Jiafei Cao, Hongrui Zhu, Zherui Qiu, Zhaxizhuoma, Yuqiang Yang, Jiaqi Peng, Xueyuan Wei, Yangkun Zhu, Jiahao Jiang, Xing Gao, Hanqing Wang, Feng Yuan, Kailin Li, Xueyue Zhu, Tai Wang, Yan Ding, Jiangmiao Pang, Jia Zeng, Jingjing Zhang, Bowen Zhou, Yao Mu, Chunhua Shen, Weinan Zhang

This paper introduces a robot control model that keeps a vision-language model’s understanding intact while adding compact future prediction, improving long-horizon manipulation and compositional generalization.

Read analysis
35Embodied Ai

In-Context World Modeling for Robotic Control

Siyin Wang, Junhao Shi, Senyu Fei, Zhaoyang Fu, Li Ji, Jingjing Gong, Xipeng Qiu

This paper teaches robots to infer how their body and environment work from a short self-generated experience window, so they can adapt to new camera angles or robot setups without retraining.

Read analysis
36Embodied Ai

Vesta: A Generalist Embodied Reasoning Model

Johan Bjorck, Zhiqi Li, Yunze Man, Jing Wang, An-Chieh Cheng, Sifei Liu, Shihao Wang, Zhiding Yu, Abhishek Badki, Stan Birchfield, Valts Blukis, Yevgen Chebotar, Siyi Chen, Sicong Leng, Yu-Cheng Chou, Tianli Ding, Boyi Li, Zhengyi Luo, Hang Su, Jonathan Tremblay, Tingwu Wang, Bowen Wen, Jimmy Wu, Xianghui Xie, Hanrong Ye, Hongxu Yin, K. R. Zentner, Liangyan Gui, Yu-Xiong Wang, Yuke Zhu, Linxi "Jim" Fan, Jan Kautz

Vesta is a robot brain that tries to do localization, navigation, spatial reasoning, and long-horizon planning in one generalist model instead of juggling multiple specialist systems.

Read analysis
37Embodied Ai

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

Yehang Zhang, Jianchong Su, Haojian Huang, Yifan Chang, Tianhao Zhou, Xinli Xu, Yingjie Xu, Yinchuan Li, Zexi Li, Ying-Cong Chen

WorldLines tests whether embodied agents can remember and reason over long-running household interactions, and pairs the benchmark with a memory system designed to keep track of changing world states.

Read analysis
39Embodied Ai

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Maggie Wang, Lars Osterberg, Stephen Tian, Ola Shorinwa, Jiajun Wu, Mac Schwager

InSight lets robot policies teach themselves new manipulation skills by breaking demonstrations into reusable primitives and then autonomously practicing, labeling, and adding missing skills to their training data.

Read analysis
40Embodied Ai

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

Yuhang Huang, Xuan Lv, Junyan Xu, Zhiyuan Yu, Jiazhao Zhang, Ruizhen Hu, Wancheng Feng, Shilong Zou, Hewen Xiao, Ziqiao Zhou, Kaiyun Huang, Zhiyu Peng, Juzhan Xu, Hang Zhao, Chenyang Zhu, Renjiao Yi, Yifei Huang, Douhui Wu, Yan Zhang, Kexu Cheng, Chunhe Song, Yunzhi Xue, Xiuhong Zhang, Leitao Guo, Yunji Chen, Bin Wu, Haibin Yu, Kai Xu

PAIWorld upgrades world models for robots by making multiple camera views agree in 3D, improving how models simulate and plan manipulation tasks.

Read analysis
42Embodied Ai

Human Universal Grasping

Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto

This paper teaches robots to grasp like humans by learning from a million egocentric human grasps and using that data to generate and retarget natural grasps in the real world.

Read analysis
43Embodied Ai

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Juncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai, Kewei Zhang, Ye Huang, Bo Liang, Shukai Gong, Jiankai Tu, Xiaotian Tang, Jiaxin Li, Kaiqi Chen, Duomin Wang, Yuqi Wang, Bingyi Kang, Eric Huang, Zhiyang Dou, Zhen Dong, Enze Xie, Wojciech Matusik, Tat-Seng Chua, Daquan Zhou

This study suggests a surprising shortcut for embodied AI: pretrain on cheap, abundant human egocentric video, then fine-tune on a small amount of robot data for the best real-world performance.

Read analysis
44Embodied Ai

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao

Embodied-R1.5 is a large embodied AI model that learns planning, correction, and grounding together, then proves it can transfer from benchmarks to real-world robot tasks.

Read analysis
45Embodied Ai

GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

Tianyi Xie, Haotian Zhang, Jinhyung Park, Zi Wang, Bowen Wen, Jiefeng Li, Xueting Li, Qingwei Ben, Haoyang Weng, Yufei Ye, David Minor, Tingwu Wang, Chenfanfu Jiang, Sanja Fidler, Jan Kautz, Linxi Fan, Yuke Zhu, Zhengyi Luo, Umar Iqbal, Ye Yuan

GRAIL uses 3D scenes plus video-model priors to generate large-scale humanoid interaction data in simulation, then turns that synthetic data into real robot skills like picking up objects and climbing stairs.

Read analysis
47Embodied Ai

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Bai, Jingren Zhou, Jiazhao Zhang, Haoqi Yuan, Gengze Zhou, Hang Yin, Ye Wang, Yiyang Huang, Zixing Lei, Wujian Peng, Delin Chen, Yingming Zheng, Jingyang Fan, Xianwei Zhuang, Xin Zhou, Haoyang Li, Anzhe Chen, Tong Zhang, Xuejing Liu, Yuchong Sun, Ruizhe Chen, Zhaohai Li, Chenxu Lü, Zhibo Yang, Tao Yu, Xionghui Chen

Qwen-VLA is a single robot brain that can see, reason, and act across different tasks and robot bodies by training one vision-language model to handle navigation, manipulation, and trajectory prediction together.

Read analysis