NTH

SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

AuthorsRuihua Han, Rui Gao, Zhe Liu, Xinyi Wang, Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao

August 27, 2026 2 min read
Watch on YouTube
The one-line take

SRL-MPC blends reinforcement learning with safety-constrained model predictive control to help differently shaped robots navigate dense crowds more safely and efficiently.

Key results

92.0%
SRL-MPC success at 25 robots

Episode success rate in the densest randomized polygon scenario.

21.0%
SARL success at 25 robots

Strongest external baseline at the same crowd density.

100.0%
Perception-noise robustness

Success rate with Gaussian neighbor-position noise of 0.02 m.

99.0%
Action-delay robustness

Success rate with a one-step, 0.1 s action delay.

10.34 ms
Controller time at 25 robots

Mean per-robot, per-step SRL-MPC computation time.

What the paper found

SRL-MPC combines model predictive control with reinforcement learning for collision-free navigation among dense crowds of robots and obstacles with arbitrary polygonal shapes. Its core representation uses geometric separation features, or GSFs, computed through support-function transformations, then embeds them in degree-2 high-order control barrier function constraints. Online planning alternates a GEOS-based geometric update with a convex local HOCBF-MPC solve; instead of replacing the planner, a PPO policy observes neighboring GSFs and continuously adjusts the path-tracking weight, control-effort weight, and safety distance. Training in IR-SIM uses 15 robots inside a 10 m × 10 m workspace, while evaluation tests 100 episodes at 10, 15, 20, and 25 robots. At the highest density, SRL-MPC achieves 92.0% episode success, compared with 21.0% for the strongest external baseline, SARL, and improves over the handcrafted adaptation rule by 17.0 percentage points. The method maintains 100.0% success under 0.02 m Gaussian perception noise and 99.0% success with a one-step, 0.1 s action delay. Its main tradeoff is computation: per-robot controller time reaches 10.34 ms at 25 robots, substantially heavier than reactive methods but still compatible with real-time control. Ablations show that both the HOCBF residual and learned continuous parameter adaptation are essential for dense polygonal crowds.

Original abstract

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability. Project website: https://hanruihua.github.io/srl_mpc_project/

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis