SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control
AuthorsRuihua Han, Rui Gao, Zhe Liu, Xinyi Wang, Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao
Resources
SRL-MPC blends reinforcement learning with safety-constrained model predictive control to help differently shaped robots navigate dense crowds more safely and efficiently.
Key results
Episode success rate in the densest randomized polygon scenario.
Strongest external baseline at the same crowd density.
Success rate with Gaussian neighbor-position noise of 0.02 m.
Success rate with a one-step, 0.1 s action delay.
Mean per-robot, per-step SRL-MPC computation time.
What the paper found
SRL-MPC combines model predictive control with reinforcement learning for collision-free navigation among dense crowds of robots and obstacles with arbitrary polygonal shapes. Its core representation uses geometric separation features, or GSFs, computed through support-function transformations, then embeds them in degree-2 high-order control barrier function constraints. Online planning alternates a GEOS-based geometric update with a convex local HOCBF-MPC solve; instead of replacing the planner, a PPO policy observes neighboring GSFs and continuously adjusts the path-tracking weight, control-effort weight, and safety distance. Training in IR-SIM uses 15 robots inside a 10 m × 10 m workspace, while evaluation tests 100 episodes at 10, 15, 20, and 25 robots. At the highest density, SRL-MPC achieves 92.0% episode success, compared with 21.0% for the strongest external baseline, SARL, and improves over the handcrafted adaptation rule by 17.0 percentage points. The method maintains 100.0% success under 0.02 m Gaussian perception noise and 99.0% success with a one-step, 0.1 s action delay. Its main tradeoff is computation: per-robot controller time reaches 10.34 ms at 25 robots, substantially heavier than reactive methods but still compatible with real-time control. Ablations show that both the HOCBF residual and learned continuous parameter adaptation are essential for dense polygonal crowds.
Original abstract
Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability. Project website: https://hanruihua.github.io/srl_mpc_project/
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.