LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation
AuthorsJin Lou, Zhiyuan Jing, Andong Chen, Xupeng Wang, Yuan Xu, Yuexuan Li, Xingdong Zhu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Ying Chu, Jingxuan Zhu, Jingyi Li, Liangliang Chen, Jinyan Liu, Zhiqi Song, Jidong Zhang, Hongming Li, Yuchen Zhu
Resources
LM-X helps generalist robots act more reliably by predicting what task stage they are in, what event comes next, and how uncertain their movements are.
Key results
Approximate size of the LM-X model built on Cosmos-Reason2-2B.
Hours of heterogeneous real-robot trajectories used for pretraining.
Hours of failed policy rollouts included to supervise off-nominal states.
Mean success improvement over the action-only backbone in five RoboTwin2.0 tasks.
LM-X mean success across 50 randomized-hard tasks, versus 55.4% for GR00T N1.7.
Mean success across seven real-robot tasks, versus 50.7% for GR00T N1.7.
What the paper found
LM-X is a generalist vision–language–action policy designed to expose its internal control state instead of returning actions as a black box. Built on the Cosmos-Reason2-2B backbone, the approximately 6B-parameter model jointly predicts three signals: Return-to-Go, or RTG, estimates task progress; Event-to-Go, or ETG, predicts the next semantic transition as a 60-step action chunk; and a heteroscedastic flow variance estimates local action reliability inside the action expert. RTG conditions ETG, and both condition a 30-step fine-grained action flow, making the explanations part of control rather than a post-hoc monitor. Training uses more than 20,000 hours of real-robot trajectories, including over 1,000 hours of failed rollouts that expose missed grasps, hesitation, and regression. In a five-task RoboTwin2.0 pretraining gate, the full design improves mean success by 16.0 percentage points over the action-only backbone. After pretraining, LM-X reaches 74.1% across 50 randomized-hard RoboTwin2.0 tasks, compared with 55.4% for NVIDIA’s GR00T N1.7, and achieves 68.6% versus 50.7% across seven real-robot tasks. RTG decreases during visible regressions, while variance spikes during oscillation and hesitation, although the paper does not establish calibrated failure detection or closed-loop recovery.
Original abstract
Generalist vision--language--action (VLA) policies learn long-horizon behavior mainly through short-horizon action prediction and reveal little beyond sampled commands. This creates two coupled bottlenecks: a single action target must implicitly absorb task progress, intermediate intent, and local reliability, while these control states remain hidden during execution. Inspired by functional principles of biological sensorimotor control, we introduce LM-X , which organizes prediction across task, event, and motor scales without claiming anatomical correspondence. Three explicitly supervised signals are emitted online and directly condition action generation: return-to-go (RTG) measures visible task progress, event-to-go (ETG) identifies the next semantic transition, and heteroscedastic action flow estimates local reliability through propagated variance. Explanation is therefore intrinsic to control rather than generated post hoc. Before a costly 20-day pretraining run on 64 NVIDIA B200 GPUs, a controlled five-task pretraining gate verifies the design: the complete model improves success by 16.0 points over the action-only backbone and by 10.8 points over the strongest single-head variant. We then train LM-X on more than 20,000 hours of real-robot trajectories, including over 1,000 hours of failed policy rollouts. LM-X achieves 74.1\% across 50 randomized-hard RoboTwin2.0 tasks versus 55.4\% for GR00T N1.7, and 68.6\% versus 50.7\% across seven real-robot tasks. RTG tracks semantic progress and visible regression, while variance rises during hesitation and oscillatory control. These results show that explicit multi-timescale predictive state can strengthen control while exposing interpretable internal estimates.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.