Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
AuthorsZijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong
Resources
This paper finds that, for RL post-training of LLMs, most of the benefit can come from training just one transformer layer—often a middle layer—rather than updating the whole model.
Key results
Qwen3-1.7B-Base single-layer GRPO on NuminaMath-CoT
Qwen3-8B-Base single-layer GRPO on NuminaMath-CoT
NuminaMath-CoT vs. DeepScaleR layer-ranking consistency
NuminaMath-CoT vs. DeepCoder layer-ranking consistency
Top 10 layers trained on Qwen3-8B-Base
Majority vote over 7 layer-trained Qwen3-1.7B models
What the paper found
This paper from the University of Minnesota, Peking University, and Amazon asks whether RL post-training in LLMs is distributed across all transformer layers or concentrated in a few. Using single-layer training with GRPO, Dr. GRPO, and GiGPO on seven models from the Qwen3 and Qwen2.5 families, including DeepSeek-Distilled-Qwen-7B, the authors define layer contribution as the fraction of full-RL gain recovered by training one layer in isolation. They find a sharply non-uniform pattern: the best layer can recover 1.14 of full-parameter RL gain on Qwen3-1.7B-Base, 1.06 on Qwen3-4B-Base, 1.07 on Qwen3-8B-Base, 1.01 on Qwen2.5-Math-1.5B, 1.02 on Qwen2.5-1.5B-Instruct, 1.01 on Qwen2.5-3B-Instruct, and 1.05 on DeepSeek-Distilled-Qwen-7B, while the weakest layers fall as low as 0.28 and even −0.51. High-contribution layers consistently cluster in the middle of the stack, and the ranking is stable across datasets and tasks, with Spearman ρ = 0.76 between NuminaMath-CoT and DeepScaleR, and 0.59 between NuminaMath-CoT and DeepCoder. Guided by this structure, layer-aware training improves Qwen3 math performance beyond full RL: on Qwen3-8B-Base, training only the top 10 layers reaches 69.1 versus 66.4 for full RL, and on Qwen3-1.7B-Base the best-guided strategy reaches 53.7 versus 50.8. The work also shows that layer-trained models solve complementary problems: majority voting over 7 layer-specialized Qwen3-1.7B models reaches 33.6% on OlympiadBench, beating the full-RL baseline at 26.9% and self-consistency voting at 31.3%.
Original abstract
Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers. Existing approaches typically update all model parameters uniformly, implicitly assuming that every layer contributes similarly to the gains obtained during RL post-training. In this work, we challenge this assumption through a systematic layer-wise study of RL training. Surprisingly, we find that training a single transformer layer can recover most of the gains achieved by full-parameter RL training, and in some cases even surpass it. To quantify this phenomenon, we introduce the quantity layer contribution, which measures the fraction of full RL improvement recovered by training a layer in isolation. Across seven models spanning two model families (Qwen3, Qwen2.5), three RL algorithms (GRPO, GiGPO, Dr. GRPO), and multiple task domains including mathematical reasoning, code generation, and agentic decision-making, we observe a remarkably stable pattern: RL gains are highly concentrated in a small subset of, and in many cases even a single, transformer layers. More strikingly, the same structural pattern consistently emerges: high-contribution layers concentrate in the middle of the transformer stack, while layers near the input and output ends contribute substantially less. The resulting layer rankings remain strongly correlated across datasets, tasks, model families, and RL algorithms.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.