NTH

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

AuthorsYifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao

June 15, 2026 2 min read
Watch on YouTube
The one-line take

Embodied-R1.5 is a large embodied AI model that learns planning, correction, and grounding together, then proves it can transfer from benchmarks to real-world robot tasks.

Key results

8B
model scale

Embodied-R1.5 parameter count

15B
training corpus

Total token scale of the embodied data system

34
datasets

Number of datasets in the unified training corpus

24
embodied VLM benchmarks

Benchmarks used for embodied VLM evaluation

16
SOTA benchmarks

Number of embodied VLM benchmarks with state-of-the-art results

70.4%
main benchmark average

Average score across the 21 main accuracy-based embodied benchmarks

What the paper found

Embodied-R1.5, developed by Tianjin University with Tencent Hunyuan as a collaborator, is a unified 8B-parameter Embodied Foundation Model that merges embodied cognition, task planning and correction, and embodied pointing into one architecture for physical intelligence. The paper’s main technical contribution is a 15B-token training system built from 34 datasets plus three automated construction pipelines: ER1.5-Spatial for tabletop 3D scene annotation, ER1.5-Correction for failure-aware planning and recovery, and ER1.5-Pointing for referring expression grounding, functional affordance grounding, and visual trace generation. Training uses a two-stage recipe—supervised fine-tuning followed by reinforced fine-tuning—with a multi-task balanced RL method that combines difficulty-aware filtering, dynamic filtering, and global batch reward normalization. On evaluation, Embodied-R1.5 reaches state-of-the-art on 16 of 24 embodied VLM benchmarks and averages 70.4% across the 21 main accuracy benchmarks, outperforming Gemini-Robotics-ER-1.5 by 17.0% and GPT-5.4 by 21.7%. The strongest gains appear in pointing and location, where it averages 72.8% and leads on all 9 pointing benchmarks, including 82.9% on Part-Afford and 80.0% on RoboAfford. The model can also be fine-tuned into Embodied-R1.5-VLA with only a small amount of action data, reaching 92.4% on SimplerEnv Google Robot Visual Matching and 97.3% on LIBERO without action pretraining. In zero-shot real-robot tests, the PGC Planner-Grounder-Corrector loop enables fully autonomous execution on tasks such as tool affordance, cup disassembly, and door opening, with 100% success on pick-and-place and tool affordance.

Original abstract

We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $π_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis