NTH

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

AuthorsYongjin Yang, Jiarui Liu, Yinghui He, Lechen Zhang, Bernhard Schölkopf, Zhijing Jin

July 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes an automated training curriculum for reasoning models that picks the next domain based not just on where the model is learning fastest, but on where an update will help other domains the most.

Key results

1.8
Macro accuracy gain on Qwen3-1.7B

TAC improves over the strongest baseline on the six-domain benchmark suite.

1.6
Macro accuracy gain on Llama3.2-3B

TAC improves over the strongest baseline on the six-domain benchmark suite.

2.8
Best improvement over learnability-only

TAC outperforms the learnability-only bandit by up to this many points.

4096
Gradient projection dimension

Transferability is computed from a 4096-dimensional Rademacher JL sketch.

0.9%
Wall-clock overhead

Transferability adds 1.06 s to a 115 s training step.

What the paper found

Transfer-Aware Curriculum, or TAC, is a multi-armed-bandit curriculum for multi-domain reinforcement learning with verifiable rewards that adds a cross-domain transferability signal to standard learnability-based sampling. Built and evaluated in the GURU six-domain reasoning suite on Qwen3-1.7B-Base and Llama3.2-3B-Instruct, TAC reuses gradients already produced by GRPO, sketches the last 4 transformer layers through a 4096-dimensional Rademacher projection, and scores domains by cosine alignment of their projected-gradient EMAs, then mixes that with normalized GRPO advantage using β = 0.2. The key result is that transferability and learnability favor different domains: TAC down-weights locally learnable but weakly transferable domains such as math and codegen, reallocates sampling toward more transferable domains like table and logic, and achieves the best macro-averaged accuracy on both backbones, beating the strongest baseline by 1.8 points on Qwen3 and 1.6 points on Llama3.2, with gains as high as 2.8 points over learnability-only sampling. The method is practical: the transfer signal adds only 1.06 seconds to a 115-second training step, or 0.9% wall-clock overhead, and the same pattern holds under an imbalanced training mix and across Qwen3-0.6B to Qwen3-4B, where TAC remains consistently better than random and SEC.

Original abstract

Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science. However, the training curriculum (how often each domain is sampled) is typically fixed or hand-tuned, even though reasoning skills transfer unevenly across domains. Existing learnability-based curricula adapt to where the policy is currently improving, but are blind to whether a gradient step on the selected domain benefits the remaining domains. In this paper, we propose Transfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose updates broadly benefit the rest of the training suite. TAC repurposes signals already produced by RL training: per-domain advantages capture local learnability, and projected gradients, taken from the GRPO step being computed, estimate cross-domain transferability via gradient-geometry alignment, at negligible cost (<1% wall-clock overhead). Across a six-domain reasoning suite, TAC achieves the best macro-averaged accuracy on both Qwen3-1.7B and Llama3.2-3B, outperforming proportional random sampling, a hand-designed schedule, and a learnability-only bandit, and improving over the last of these by up to 2.8 points (10% relative). Ablations show performance degrades sharply when the transferability term is removed, and TAC remains robust on imbalanced training mixtures where learnability-only curricula over-commit to dominant domains. Our findings establish cross-domain transferability as a key signal for curriculum design in multi-domain RLVR.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →