Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding
AuthorsXuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang, Sean Du, Chujie Zheng, Fei Huang, Dayiheng Liu, Gao Huang, Jingren Zhou
Resources
This paper shows that for some aligned LLMs, the best next-token predictions may come from a near-final layer rather than the last one, and it uses that insight to improve reasoning with almost no extra cost.
Key results
Share of tokens where the backward scan selected a non-final entropy valley layer
Share of generated tokens whose argmax changed under Confident Decoding
End-to-end wall-clock latency increase reported as under 2% per token
Absolute improvement from last-layer decoding to Confident Decoding
Absolute improvement on code generation benchmark
Absolute improvement on Level 4 Omni-MATH
What the paper found
The paper from Alibaba’s Qwen Team, with coauthors from Tsinghua University and Nanyang Technological University, argues that “deeper” transformer layers are not always the best decoding point in aligned LLMs. By analyzing Qwen3.5 models, it identifies a recurring Guess–Refine–Perturb pattern: early layers guess, middle layers sharpen reasoning, and the final layers can perturb predictions toward generic or safety-biased tokens. The proposed Confident Decoding method is training-free and drop-in: at each token step it computes intermediate logits over a near-final window, then backtracks to the first local minimum in predictive entropy, which the authors call the entropy valley. On Qwen3.5-35B-A3B, the method selects a non-final layer for 11.5% of tokens, but only 2.47% of all generated tokens actually change their argmax, while end-to-end latency stays under 2% with zero extra KV-cache memory. Across dense and MoE models, it improves GPQA-Diamond, HLE, LiveCodeBench v6, and Omni-MATH, with especially large reasoning gains such as +6.5 on GPQA-Diamond and +10.1 on LiveCodeBench v6 for Qwen3.5-27B, and +22.4 on the hardest Omni-MATH tier for gpt-oss-20b. The paper also shows the effect is stronger in instruction-tuned models than base models, supporting its claim that post-training alignment introduces an “alignment tax” that Confident Decoding can selectively bypass without sacrificing safety on Air-Bench-2024 or writing quality on WritingBench.
Original abstract
Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred tokens. We introduce Confident Decoding, a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy-guided conservative backward search. We further provide a theoretical formulation of layer selection as an optimal stopping problem, showing that under bounded projection noise and dominant late-stage alignment perturbation, our search rule filters perturbation while bounding the loss relative to the oracle refinement layer. Experiments across dense and Mixture-of-Experts LLMs demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE, with zero memory overhead and less than 2% latency increase. These results suggest dynamically bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.