Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
AuthorsMartin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Brüdigam
Resources
The work uses simulation-based control demonstrations to teach quadrupeds and humanoids complex locomotion-and-manipulation skills with sparse rewards and minimal manual tuning.
Key results
Initial proportion of SMPC transitions mixed into the online replay buffer.
Demonstration samples generated per hour through GPU-parallelized simulation.
SMPC samples required for convergence on the hardest task.
GPU hours needed to collect the 4M-sample bootstrap dataset.
What the paper found
This paper presents an offline-to-online reinforcement-learning pipeline for whole-body loco-manipulation that replaces manual dense-reward engineering with Sample-based Model Predictive Control, or SMPC, demonstrations generated entirely in simulation. Using MuJoCo Warp and a modified FastTD3 algorithm, the system initially mixes 50% expert transitions into the replay buffer, then phases them out after the policy begins succeeding, while optimizing only sparse task rewards. A frozen ReLIC whole-body stabilization controller converts high-level velocity, arm, torso-height, and pitch commands into dynamically stable joint targets, enabling neural policies to improve beyond their SMPC teacher rather than merely imitate it. GPU-vectorized SMPC produces 1M samples per hour, allowing the most difficult task to bootstrap from 4M samples collected within 4 GPU hours. The method reaches near-perfect simulated success and transfers five tasks—including reaching, box pushing, tire uprighting, tire rolling, and humanoid box pushing—to hardware on an arm-equipped Boston Dynamics Spot and a Unitree G1. Learned policies complete some tasks more than 50% faster than SMPC and reduce task-duration variability by 11–45%. Ablations show that larger datasets are needed for more complex coordination, while multimodal demonstrations can prevent learning; filtering SMPC data to a single behavioral mode is therefore critical.
Original abstract
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets. Because this data solves the fundamental exploration problem, we can train an off-policy RL agent using purely sparse task rewards, drastically reducing the time required to learn new skills and eliminating the need for manual tuning. Integrating this high-level agent with a low-level dynamic stability controller yields more optimal behaviors that strictly align with true task objectives, ultimately allowing the learned policies to surpass the original optimal control teacher. We validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies, including an arm-equipped Spot quadruped and a G1 humanoid.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.