NTH

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

AuthorsZijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong

July 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper finds that, for RL post-training of LLMs, most of the benefit can come from training just one transformer layer—often a middle layer—rather than updating the whole model.

Key results

1.14
best layer contribution

Qwen3-1.7B-Base single-layer GRPO on NuminaMath-CoT

-0.51
worst layer contribution

Qwen3-8B-Base single-layer GRPO on NuminaMath-CoT

0.76
Spearman rho math datasets

NuminaMath-CoT vs. DeepScaleR layer-ranking consistency

0.59
Spearman rho math vs code

NuminaMath-CoT vs. DeepCoder layer-ranking consistency

69.1
Qwen3-8B guided math

Top 10 layers trained on Qwen3-8B-Base

33.6
OlympiadBench voting

Majority vote over 7 layer-trained Qwen3-1.7B models

What the paper found

This paper from the University of Minnesota, Peking University, and Amazon asks whether RL post-training in LLMs is distributed across all transformer layers or concentrated in a few. Using single-layer training with GRPO, Dr. GRPO, and GiGPO on seven models from the Qwen3 and Qwen2.5 families, including DeepSeek-Distilled-Qwen-7B, the authors define layer contribution as the fraction of full-RL gain recovered by training one layer in isolation. They find a sharply non-uniform pattern: the best layer can recover 1.14 of full-parameter RL gain on Qwen3-1.7B-Base, 1.06 on Qwen3-4B-Base, 1.07 on Qwen3-8B-Base, 1.01 on Qwen2.5-Math-1.5B, 1.02 on Qwen2.5-1.5B-Instruct, 1.01 on Qwen2.5-3B-Instruct, and 1.05 on DeepSeek-Distilled-Qwen-7B, while the weakest layers fall as low as 0.28 and even −0.51. High-contribution layers consistently cluster in the middle of the stack, and the ranking is stable across datasets and tasks, with Spearman ρ = 0.76 between NuminaMath-CoT and DeepScaleR, and 0.59 between NuminaMath-CoT and DeepCoder. Guided by this structure, layer-aware training improves Qwen3 math performance beyond full RL: on Qwen3-8B-Base, training only the top 10 layers reaches 69.1 versus 66.4 for full RL, and on Qwen3-1.7B-Base the best-guided strategy reaches 53.7 versus 50.8. The work also shows that layer-trained models solve complementary problems: majority voting over 7 layer-specialized Qwen3-1.7B models reaches 33.6% on OlympiadBench, beating the full-RL baseline at 26.9% and self-consistency voting at 31.3%.

Original abstract

Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers. Existing approaches typically update all model parameters uniformly, implicitly assuming that every layer contributes similarly to the gains obtained during RL post-training. In this work, we challenge this assumption through a systematic layer-wise study of RL training. Surprisingly, we find that training a single transformer layer can recover most of the gains achieved by full-parameter RL training, and in some cases even surpass it. To quantify this phenomenon, we introduce the quantity layer contribution, which measures the fraction of full RL improvement recovered by training a layer in isolation. Across seven models spanning two model families (Qwen3, Qwen2.5), three RL algorithms (GRPO, GiGPO, Dr. GRPO), and multiple task domains including mathematical reasoning, code generation, and agentic decision-making, we observe a remarkably stable pattern: RL gains are highly concentrated in a small subset of, and in many cases even a single, transformer layers. More strikingly, the same structural pattern consistently emerges: high-contribution layers concentrate in the middle of the transformer stack, while layers near the input and output ends contribute substantially less. The resulting layer rankings remain strongly correlated across datasets, tasks, model families, and RL algorithms.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →