Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
AuthorsMichael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville
AffiliationsMila, Université de Montréal · Allen Institute for AI · University of Washington · Trillium Labs · Canada CIFAR AI Chair
Resources
NGU helps RL-trained LLMs stop over-practicing easy problems and spend more effort solving the hard ones.
Key results
Probability of continuing to resample an all-incorrect prompt in the main NGU experiment.
Number of Deepscaler math problems used to train Qwen 3 4B-base.
Average score across AIME 2025 and BRUMO November 2025, versus 24.5 for the fixed K=64 baseline.
Pass@1 improvement over the initial model for Deepscaler hard problems with NGU pNGU=0.875.
Extra-hard pass@1 for NGU with asynchronous RL and positive anchoring.
Approximate fraction of Manufactoria tests learned by standard GRPO without fully solving any problem.
What the paper found
This paper identifies a Matthew Effect in reinforcement learning for large language models: RL improves tasks in proportion to the model’s initial competence, making easy problems easier while leaving hard problems under-trained. The authors attribute this less to insufficient samples than to poor compute allocation in fixed-group methods such as GRPO, which repeatedly spends rollouts on already-solvable prompts. Their solution, Never Give Up, uses asynchronous RL to resample an all-incorrect prompt with probability pNGU=0.95, retaining recent completions and filtering stale off-policy negatives; easy prompts exit quickly, while difficult prompts receive additional search. On a 10k subset of Deepscaler, Qwen 3 4B-base with NGU reached an average pass@1 of 26.5 across AIME 2025 and BRUMO November 2025, compared with 24.5 for the strongest reported fixed K=64 baseline, while improving the hard subset by 4.3 points. The method also requires asynchronous infrastructure and performs best when completions older than T=4 steps are excluded, with “anchoring the positives” used to rescale remaining advantages. On Manufactoria, standard GRPO learned roughly 80% of individual coding tests but never fully solved a problem; NGU continued improving on the hardest tests and reached 87.9% all-tests pass@1 in the GSM8k ablation’s extra-hard setting. The findings refine lessons from DeepSeek-R1 and related Qwen post-training: scaling RL requires adaptive signal efficiency, not merely more samples per prompt.
Original abstract
We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvements on easy problems that an LLM is already good at solving, but small improvements on hard problems. We call this the Matthew Effect in RL for LLMs, after the phenomenon of cumulative advantage from economics and network science summarized as "the rich get richer". The naive explanation is that hard problems require more compute to find a solution. We argue that modern RL methods are exacerbating the issue by wasting too much compute on easy problems and instead should dynamically reallocate how they use compute. We introduce Never Give Up (NGU), a simple adaptive sampling method that keeps generating samples for a problem until one is correct. By leveraging asynchronous RL, this naturally uses fewer samples to filter out easy problems and allocates more compute to solving harder problems. We investigate the design choices that affect NGU, such as off-policy robustness, and develop a set of best practices. On the math benchmark Deepscaler, NGU improves performance per compute, especially on harder problems. On a recent coding task, Manufactoria, standard GRPO with a per-test reward fails to fully solve problems that have a range of easy and difficult tests. NGU iteratively improves, solving harder and harder tests, until it learns to fully solve coding problems.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.