DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
AuthorsGuochao Jiang, Jingyi Song, Guofeng Quan, Chuzhan Hao, Guohua Liu, Yuewei Zhang
Resources
DVAO is a new reinforcement learning method that adapts reward weighting on the fly to stabilize multi-objective training for LLMs and improve performance on reasoning and tool-use tasks.
Key results
Rollout group size used for variance estimation and DVAO weighting
Training batch size per rollout batch
Number of optimization steps used in experiments
Constant AdamW learning rate
DVAO average accuracy on mathematical reasoning benchmarks
DVAO average format compliance on BFCL-v4
What the paper found
DVAO, or Dynamic Variance-adaptive Advantage Optimization, is an Alibaba Cloud proposal for multi-reward reinforcement learning in large language model alignment that targets the failure modes of standard GRPO scalarization. The paper argues that Reward Combination can inflate squared advantage magnitudes and destabilize policy gradients, while Advantage Combination with static weights ignores cross-objective correlations. DVAO replaces fixed weights with variance-adaptive weights proportional to each reward’s empirical rollout-group standard deviation, so high-variance objectives receive more learning signal and noisy low-variance objectives are suppressed. The authors prove two key properties: DVAO bounds advantage magnitude relative to raw reward combination and introduces an implicit cross-objective regularization effect, with sensitivity depending on the combined rollout performance rather than isolated objective scores. Experiments use Qwen3-4B-Base and Qwen3-8B-Base on mathematical reasoning benchmarks AIME-2024, AIME-2025, MATH500, OlympiadBench, and AMC23, plus Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on BFCL-v4 tool-use. DVAO reaches the strongest average trade-offs, such as 47.49 accuracy and 99.92 length compliance on Qwen3-8B-Base, and 63.00 accuracy with 79.21 average format compliance on Qwen2.5-7B-Instruct, while dominating the Pareto frontier across weight sweeps. Training uses DAPO-MATH-17K, ToolRL data, group size 16, batch size 128, 500 steps, AdamW at 1e-6, and 8,192-token generation.
Original abstract
Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative Policy Optimization offers an efficient, value-model-free alternative to Proximal Policy Optimization, adapting it to real-world multi-reward settings remains challenging. Standard scalarization practices, such as Reward Combination and Advantage Combination, suffer from significant drawbacks: Reward Combination frequently generates advantages with excessively large squared magnitudes that lead to training instability, while Advantage Combination relies on static hyperparameters and ignores cross-objective correlations. To address these limitations, we propose Dynamic Variance-adaptive Advantage Optimization (DVAO), which dynamically adjusts combination weights based on the empirical reward variance of each objective within a rollout group, effectively up-weighting objectives with a stronger learning signal while suppressing noisy ones. We mathematically prove that DVAO maintains bounded advantage magnitudes for stable training and introduces a self-adaptive cross-objective regularization mechanism. Extensive experiments on mathematical reasoning and tool-use benchmarks using Qwen3 and Qwen2.5 models demonstrate that DVAO significantly outperforms baseline methods, achieving a superior multi-objective Pareto frontier and robust training stability.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.