MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
AuthorsQiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
AffiliationsXspark AI · The Hong Kong University of Science and Technology (Guangzhou) · Tsinghua University · The University of Hong Kong
Resources
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Key results
Success rate on the EBench mobile-manipulation benchmark.
Task-weighted success across RoboCasa365 evaluation splits.
Mean success using the shared policy on fixed-base LIBERO tasks.
Success under perturbations without additional perturbation training.
Mean success across five long-horizon mobile-manipulation tasks.
Full-model success on RoboCasa365 composite-seen tasks, compared with 29.2% for velocity prediction.
What the paper found
MM-ABC is a generalist mobile-manipulation foundation model designed to coordinate a robot’s arms and mobile base as one dynamically changing workspace. Its Seeing component uses sparse DeepStack conditioning from intermediate and final layers of the Qwen3-VL-4B backbone; Coordinating is handled by MM-APT, a dual-stream transformer with masked joint attention, asymmetric near–far visibility, and clean-action x-prediction instead of velocity prediction; Imagining adds a training-only branch that predicts future VGGT-Ω geometric features, strengthening spatial and manipulation-intent representations without adding inference-time cost. The model is pretrained on more than 5,000 hours spanning 12 datasets and 17 embodiments, then evaluated across mobile and fixed-base settings. It reaches 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, and 82.8% on LIBERO-Plus without perturbation training. On five real-world tasks, it averages 83% success, compared with 71% for π0.5 and 54% for StarVLA-GR00T. Ablations show that the architectural choices matter: on RoboCasa365 composite-seen tasks, the full model scores 32.8%, versus 29.2% with velocity prediction and 24.9% without future supervision. Compared with systems such as OpenVLA and GR00T, MM-ABC emphasizes cross-stream collaboration and anticipatory geometry rather than simply enlarging a single action decoder.
Original abstract
Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.
Agent as Policy for Robotic Manipulation
Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang
A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.