Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
AuthorsJunhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji, Jingjing Gong, Xipeng Qiu
Resources
This paper says robots should first learn how to move from cheap unlabeled experience, and only then learn what humans want them to do from scarce expert demonstrations.
Key results
TAP-20k success rate on SIMPLER
Baseline on SIMPLER with same stage-2 data
Task-agnostic pretraining episodes used in the best SIMPLER setting
Expert trajectories used for simulation fine-tuning
Hours of autonomous random play for real-world pretraining
Expert demonstrations per task in WidowX experiments
What the paper found
This Fudan University paper introduces Task-Agnostic Pretraining, or TAP, for vision-language-action models by separating “how to move” from “what to do.” Instead of treating every robot trajectory as a language-labeled expert demonstration, TAP first pretrains a Qwen2.5-VL 3B backbone with a SigLIP visual encoder on cheap unlabeled interaction data, including discarded Bridge trajectories and autonomous random play, using a self-supervised inverse dynamics objective that predicts the action between two observations. The second stage then fine-tunes on a small expert set with standard behavior cloning. On the SIMPLER benchmark, TAP-20k reaches 33.32% Avg-All success, outperforming Standard BC at 23.15% by 10% absolute and approaching internet-scale baselines pretrained on Open X-Embodiment, while using only 5k expert trajectories in stage 2. In real-world WidowX 250s experiments with only 200 expert demonstrations, TAP is pretrained on 30 hours of self-exploration and achieves strong robustness: under background texture shift it scores 25% on “put the carrot on the plate” versus 10% for NORA, and under viewpoint variation it still reaches 15% and 25% where NORA collapses to 0% on both tasks. The ablation studies show that scaling stage-1 exposure from 20k to 100k steps raises Avg-All from 18% to above 30%, and Grad-CAM visualizations confirm that inverse dynamics pretraining concentrates attention on manipulable objects before any language instruction is introduced.
Original abstract
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap, unlabeled interaction data -- including discarded off-task trajectories and autonomous robot play -- via a self-supervised Inverse Dynamics objective. A lightweight second stage then grounds these priors in language using minimal expert data. On the SIMPLER benchmark, TAP matches models trained on over 1M expert trajectories while using orders of magnitude less labeled data, yielding a 10% absolute gain over standard behavior cloning. On a real-world WidowX platform, TAP retains 25% success under camera perturbations where internet-scale baselines collapse to 0%, demonstrating that task-agnostic pretraining produces robust, transferable physical representations and offers a scalable path forward for Embodied AI.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.