WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
AuthorsZihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan, Fei Ni, Jinyi Liu, Wei Wei, Jianrong Wang, Yan Zheng, Jianye Hao
Resources
WarpSAC adapts exploration and exploitation strategies to the amount of available simulation data, substantially improving scalable reinforcement learning and robot training.
Key results
Normalized score–step AUC improvement over FlashSAC across nine CPU-scale environments.
Normalized score–step AUC improvement over FlashSAC across fourteen GPU-parallel environments.
Starting success rate on UnitreeG1TransportBox-v1 before WarpSAC-A.
Success rate achieved by WarpSAC-A on UnitreeG1TransportBox-v1.
Wall-clock reduction versus FlashSAC in Unitree G1 sim-to-real training.
What the paper found
WarpSAC rethinks scalable off-policy reinforcement learning by treating exploration and exploitation as data-regime problems rather than fixed algorithmic requirements. Built on Soft Actor-Critic and FlashSAC, it isolates three mechanisms: Sample Weight Decay, which uses linear age-based replay weighting to prioritize policy-relevant transitions; parameter projection normalization; and clipped double-Q critics. Across eight benchmark families, including the DeepMind Control Suite, HumanoidBench, MyoSuite, MuJoCo Playground, IsaacLab, MJLab, and ManiSkill, the study finds that normalization and double-Q improve learning when replay coverage is narrow, but can restrict value fitting or add pessimism when thousands of GPU-parallel environments provide abundant data. WarpSAC-L therefore keeps normalization and double-Q for CPU-scale training, while WarpSAC-A removes both and uses a single critic for GPU-parallel training; Sample Weight Decay remains enabled in both. Against FlashSAC, normalized score–step AUC improves by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. On UnitreeG1TransportBox-v1, success rises from 19.8% to 96.4%, while MuJoCo Playground wall-time AUC improves by 19.1%. In a Unitree Robotics G1 sim-to-real locomotion pipeline running on an NVIDIA A800, WarpSAC reduces wall-clock training time by 36.4%, reaching deployment performance in 35 minutes. The central result is that scalable off-policy RL often becomes faster and stronger by removing conservative components when data are plentiful, not by stacking more stabilizers.
Original abstract
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.