NTH

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

AuthorsGuochao Jiang, Jingyi Song, Guofeng Quan, Chuzhan Hao, Guohua Liu, Yuewei Zhang

June 26, 2026 2 min read
Watch on YouTube
The one-line take

DVAO is a new reinforcement learning method that adapts reward weighting on the fly to stabilize multi-objective training for LLMs and improve performance on reasoning and tool-use tasks.

Key results

16
group size

Rollout group size used for variance estimation and DVAO weighting

128
prompt batch size

Training batch size per rollout batch

500
training steps

Number of optimization steps used in experiments

1e-6
learning rate

Constant AdamW learning rate

47.49
Qwen3-8B-Base average accuracy

DVAO average accuracy on mathematical reasoning benchmarks

79.21
Qwen2.5-7B-Instruct average format compliance

DVAO average format compliance on BFCL-v4

What the paper found

DVAO, or Dynamic Variance-adaptive Advantage Optimization, is an Alibaba Cloud proposal for multi-reward reinforcement learning in large language model alignment that targets the failure modes of standard GRPO scalarization. The paper argues that Reward Combination can inflate squared advantage magnitudes and destabilize policy gradients, while Advantage Combination with static weights ignores cross-objective correlations. DVAO replaces fixed weights with variance-adaptive weights proportional to each reward’s empirical rollout-group standard deviation, so high-variance objectives receive more learning signal and noisy low-variance objectives are suppressed. The authors prove two key properties: DVAO bounds advantage magnitude relative to raw reward combination and introduces an implicit cross-objective regularization effect, with sensitivity depending on the combined rollout performance rather than isolated objective scores. Experiments use Qwen3-4B-Base and Qwen3-8B-Base on mathematical reasoning benchmarks AIME-2024, AIME-2025, MATH500, OlympiadBench, and AMC23, plus Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on BFCL-v4 tool-use. DVAO reaches the strongest average trade-offs, such as 47.49 accuracy and 99.92 length compliance on Qwen3-8B-Base, and 63.00 accuracy with 79.21 average format compliance on Qwen2.5-7B-Instruct, while dominating the Pareto frontier across weight sweeps. Training uses DAPO-MATH-17K, ToolRL data, group size 16, batch size 128, 500 steps, AdamW at 1e-6, and 8,192-token generation.

Original abstract

Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative Policy Optimization offers an efficient, value-model-free alternative to Proximal Policy Optimization, adapting it to real-world multi-reward settings remains challenging. Standard scalarization practices, such as Reward Combination and Advantage Combination, suffer from significant drawbacks: Reward Combination frequently generates advantages with excessively large squared magnitudes that lead to training instability, while Advantage Combination relies on static hyperparameters and ignores cross-objective correlations. To address these limitations, we propose Dynamic Variance-adaptive Advantage Optimization (DVAO), which dynamically adjusts combination weights based on the empirical reward variance of each objective within a rollout group, effectively up-weighting objectives with a stronger learning signal while suppressing noisy ones. We mathematically prove that DVAO maintains bounded advantage magnitudes for stable training and introduces a self-adaptive cross-objective regularization mechanism. Extensive experiments on mathematical reasoning and tool-use benchmarks using Qwen3 and Qwen2.5 models demonstrate that DVAO significantly outperforms baseline methods, achieving a superior multi-objective Pareto frontier and robust training stability.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →