Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR
AuthorsYongjin Yang, Jiarui Liu, Yinghui He, Lechen Zhang, Bernhard Schölkopf, Zhijing Jin
Resources
This paper proposes an automated training curriculum for reasoning models that picks the next domain based not just on where the model is learning fastest, but on where an update will help other domains the most.
Key results
TAC improves over the strongest baseline on the six-domain benchmark suite.
TAC improves over the strongest baseline on the six-domain benchmark suite.
TAC outperforms the learnability-only bandit by up to this many points.
Transferability is computed from a 4096-dimensional Rademacher JL sketch.
Transferability adds 1.06 s to a 115 s training step.
What the paper found
Transfer-Aware Curriculum, or TAC, is a multi-armed-bandit curriculum for multi-domain reinforcement learning with verifiable rewards that adds a cross-domain transferability signal to standard learnability-based sampling. Built and evaluated in the GURU six-domain reasoning suite on Qwen3-1.7B-Base and Llama3.2-3B-Instruct, TAC reuses gradients already produced by GRPO, sketches the last 4 transformer layers through a 4096-dimensional Rademacher projection, and scores domains by cosine alignment of their projected-gradient EMAs, then mixes that with normalized GRPO advantage using β = 0.2. The key result is that transferability and learnability favor different domains: TAC down-weights locally learnable but weakly transferable domains such as math and codegen, reallocates sampling toward more transferable domains like table and logic, and achieves the best macro-averaged accuracy on both backbones, beating the strongest baseline by 1.8 points on Qwen3 and 1.6 points on Llama3.2, with gains as high as 2.8 points over learnability-only sampling. The method is practical: the transfer signal adds only 1.06 seconds to a 115-second training step, or 0.9% wall-clock overhead, and the same pattern holds under an imbalanced training mix and across Qwen3-0.6B to Qwen3-4B, where TAC remains consistently better than random and SEC.
Original abstract
Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science. However, the training curriculum (how often each domain is sampled) is typically fixed or hand-tuned, even though reasoning skills transfer unevenly across domains. Existing learnability-based curricula adapt to where the policy is currently improving, but are blind to whether a gradient step on the selected domain benefits the remaining domains. In this paper, we propose Transfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose updates broadly benefit the rest of the training suite. TAC repurposes signals already produced by RL training: per-domain advantages capture local learnability, and projected gradients, taken from the GRPO step being computed, estimate cross-domain transferability via gradient-geometry alignment, at negligible cost (<1% wall-clock overhead). Across a six-domain reasoning suite, TAC achieves the best macro-averaged accuracy on both Qwen3-1.7B and Llama3.2-3B, outperforming proportional random sampling, a hand-designed schedule, and a learnability-only bandit, and improving over the last of these by up to 2.8 points (10% relative). Ablations show performance degrades sharply when the transferability term is removed, and TAC remains robust on imbalanced training mixtures where learnability-only curricula over-commit to dominant domains. Our findings establish cross-domain transferability as a key signal for curriculum design in multi-domain RLVR.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.