NTH

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

AuthorsZihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan, Fei Ni, Jinyi Liu, Wei Wei, Jianrong Wang, Yan Zheng, Jianye Hao

August 27, 2026 2 min read
Watch on YouTube
The one-line take

WarpSAC adapts exploration and exploitation strategies to the amount of available simulation data, substantially improving scalable reinforcement learning and robot training.

Key results

4.5%
CPU-scale AUC improvement

Normalized score–step AUC improvement over FlashSAC across nine CPU-scale environments.

23.1%
GPU-parallel AUC improvement

Normalized score–step AUC improvement over FlashSAC across fourteen GPU-parallel environments.

19.8%
Unitree transport success baseline

Starting success rate on UnitreeG1TransportBox-v1 before WarpSAC-A.

96.4%
Unitree transport success

Success rate achieved by WarpSAC-A on UnitreeG1TransportBox-v1.

36.4%
G1 sim-to-real wall-time reduction

Wall-clock reduction versus FlashSAC in Unitree G1 sim-to-real training.

What the paper found

WarpSAC rethinks scalable off-policy reinforcement learning by treating exploration and exploitation as data-regime problems rather than fixed algorithmic requirements. Built on Soft Actor-Critic and FlashSAC, it isolates three mechanisms: Sample Weight Decay, which uses linear age-based replay weighting to prioritize policy-relevant transitions; parameter projection normalization; and clipped double-Q critics. Across eight benchmark families, including the DeepMind Control Suite, HumanoidBench, MyoSuite, MuJoCo Playground, IsaacLab, MJLab, and ManiSkill, the study finds that normalization and double-Q improve learning when replay coverage is narrow, but can restrict value fitting or add pessimism when thousands of GPU-parallel environments provide abundant data. WarpSAC-L therefore keeps normalization and double-Q for CPU-scale training, while WarpSAC-A removes both and uses a single critic for GPU-parallel training; Sample Weight Decay remains enabled in both. Against FlashSAC, normalized score–step AUC improves by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. On UnitreeG1TransportBox-v1, success rises from 19.8% to 96.4%, while MuJoCo Playground wall-time AUC improves by 19.1%. In a Unitree Robotics G1 sim-to-real locomotion pipeline running on an NVIDIA A800, WarpSAC reduces wall-clock training time by 36.4%, reaching deployment performance in 35 minutes. The central result is that scalable off-policy RL often becomes faster and stronger by removing conservative components when data are plentiful, not by stacking more stabilizers.

Original abstract

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →