NTH

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

AuthorsXuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang, Sean Du, Chujie Zheng, Fei Huang, Dayiheng Liu, Gao Huang, Jingren Zhou

June 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that for some aligned LLMs, the best next-token predictions may come from a near-final layer rather than the last one, and it uses that insight to improve reasoning with almost no extra cost.

Key results

11.5%
Qwen3.5-35B-A3B token selection rate

Share of tokens where the backward scan selected a non-final entropy valley layer

2.47%
Qwen3.5-35B-A3B substitution rate

Share of generated tokens whose argmax changed under Confident Decoding

2%
Latency overhead

End-to-end wall-clock latency increase reported as under 2% per token

6.5
Qwen3.5-27B GPQA-Diamond gain

Absolute improvement from last-layer decoding to Confident Decoding

10.1
Qwen3.5-27B LiveCodeBench v6 gain

Absolute improvement on code generation benchmark

22.4
gpt-oss-20b Omni-MATH hardest-tier gain

Absolute improvement on Level 4 Omni-MATH

What the paper found

The paper from Alibaba’s Qwen Team, with coauthors from Tsinghua University and Nanyang Technological University, argues that “deeper” transformer layers are not always the best decoding point in aligned LLMs. By analyzing Qwen3.5 models, it identifies a recurring Guess–Refine–Perturb pattern: early layers guess, middle layers sharpen reasoning, and the final layers can perturb predictions toward generic or safety-biased tokens. The proposed Confident Decoding method is training-free and drop-in: at each token step it computes intermediate logits over a near-final window, then backtracks to the first local minimum in predictive entropy, which the authors call the entropy valley. On Qwen3.5-35B-A3B, the method selects a non-final layer for 11.5% of tokens, but only 2.47% of all generated tokens actually change their argmax, while end-to-end latency stays under 2% with zero extra KV-cache memory. Across dense and MoE models, it improves GPQA-Diamond, HLE, LiveCodeBench v6, and Omni-MATH, with especially large reasoning gains such as +6.5 on GPQA-Diamond and +10.1 on LiveCodeBench v6 for Qwen3.5-27B, and +22.4 on the hardest Omni-MATH tier for gpt-oss-20b. The paper also shows the effect is stronger in instruction-tuned models than base models, supporting its claim that post-training alignment introduces an “alignment tax” that Confident Decoding can selectively bypass without sacrificing safety on Air-Bench-2024 or writing quality on WritingBench.

Original abstract

Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred tokens. We introduce Confident Decoding, a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy-guided conservative backward search. We further provide a theoretical formulation of layer selection as an optimal stopping problem, showing that under bounded projection noise and dominant late-stage alignment perturbation, our search rule filters perturbation while bounding the loss relative to the oracle refinement layer. Experiments across dense and Mixture-of-Experts LLMs demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE, with zero memory overhead and less than 2% latency increase. These results suggest dynamically bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis