Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Deeper is Not Always Better".
Tom: The paper introduces Confident Decoding to address the "Alignment Tax," a phenomenon where final-layer perturbations in large language models disrupt reasoning-relevant predictions,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we're looking at this paper called "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding." It tackles this idea that just making the model deeper always makes it smarter, and claims that it doesn't actually hold up when you consider how AI models are trained.
Jane: Right, so they're pointing out a problem where those final layers of a large language model can actually mess things up. They suggest that these later layers introduce some kind of noise that pushes the model away from the actual reasoning it should be doing, sometimes toward generic answers instead of complex logic.
Lu: What's really interesting here is their analysis of what they call a "Guess–Refine–Perturb" pattern in how these models process information during a forward pass. It seems like early layers make some initial guesses, middle layers fine-tune the reasoning part, and then the final layers introduce that shift.
Meng: So if this perturbation happens at the end, it means we can't just rely on whatever that last layer spits out for our most important answers anymore because it might be a weak signal.
Lalam: That makes sense when you think about how I work; I try to keep my core reasoning solid, but if the final output layer gets too noisy, it can make my complex logic seem flimsy.
Tom: Exactly, and that’s where this paper comes in with Confident Decoding. It proposes a training-free way to bypass those late-stage biases by dynamically picking a reliable layer rather than just going straight to the end.
Jane: They describe this strategy as using an entropy guided conservative backward search to pick the best near-final layer, which means it's not retraining anything or changing how we run the model at all.
Lu: It’s quite clever because they’re not truncating the transformer or messing with the forward pass structure; they are just smart about where to look within it based on certain dynamics.
Paper summary: Meng: So, how does this dynamic selection actually work in practice when you have these complex layer contributions and residual similarities? We need to know if this is just theoretical fluff.
Tom: Well, they use metrics like Relative Contribution Norm and Residual I/O Cosine Similarity to formally characterize those three phases of the process—Guess, Refine, and Perturb—as detailed in "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding."
Jane: The paper lays out a risk function, R(l), which shows that this late-stage perturbation acts differently depending on what the AI is trying to do.
Lu: They show that for safety tasks, like following simple guardrails, this noise might actually be helpful; but for complex reasoning, it seems to derail the established logic chain instead.
Tom: That distinction between "Tax vs. Guardrail" is a really important conceptual move here. It suggests that what we consider an alignment problem for one kind of task might not be the same for another.
Meng: So if we're focused on complex reasoning, this paper claims Confident Decoding secures substantial gains on difficult benchmarks like GPQA-Diamond and Omni-MATH across different architectures.
Jane: They validate this by showing that the method consistently secures substantial reasoning gains across all models they tested, which is a strong empirical result in "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding."
Lu: The authors specifically point out that this strategy aggressively neutralizes what they call the "Planning–Pragmatics Tradeoff," which they say is most destructive when dealing with fragile, low-frequency logic chains.
Tom: That sounds like it directly addresses one of the main weaknesses we see in current LLMs when they try to handle multi-step reasoning problems.
Jane: And computationally, they’re pretty good about that; the overhead is small because you’re just doing a few extra calculations, bounded by one batched unembedding of M candidate hidden states and a K-step scan.
Meng: So if the overhead is less than two percent per token latency increase, that makes it very practical for deployment in real applications without slowing everything down significantly.
Paper summary: Lalam: From my side, having a deterministic way to select the most reliable layer means I can trust the output more consistently, which is a big plus for cultural applications where reliability matters a lot.
Tom: It really comes down to this: Confident Decoding successfully isolates the generation process from those corrupted terminal distributions by finding that monotonic entropy valley during that backward scan.
Jane: The authors conclude that this dynamic selection proves we can unlock stronger reasoning behavior from models that are already aligned, even when there’s an alignment tax present in the final layers of their structure.
Lu: They default to a deterministic setting for this valley selection because it outperforms stochastic mixtures by bypassing that alignment-tax bias, suggesting a very direct path forward.
Meng: For practical implementation right now, this means we can start testing models on those challenging benchmarks and expect better performance without having to overhaul the entire model architecture or retraining it from scratch.
Tom: So to wrap up this discussion on "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding," we see a method that isolates generation from those final layer problems by smartly navigating the internal dynamics of how an AI generates text.
Jane: It’s about finding that point where the model has high confidence before any potential late-layer perturbation kicks in, which is a really direct way to improve complex reasoning capabilities.
Lu: Future work they suggest could look at training paradigms where you apply alignment penalties only to specific routing heads instead of the entire core residual stream, which would be a different approach entirely.
Meng: That sounds like a more structural change to how we design the AI itself rather than just fixing the decoding step.
Tom: Yeah, that's interesting because it moves the problem from runtime correction into model training design, but for now Confident Decoding gives us a solid tool to boost performance on existing models.
Jane: So this paper shows that depth isn't automatically better when alignment is involved; we have to be smarter about where we look in the model's layers.
Conclusion: Tom: So we've been looking at how these models can get stuck in traps where their final layers mess up the actual reasoning, and this paper is called "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding."
Jane: Yeah, it’s basically about finding a way to bypass that noise at the end of the generation process without having to retrain everything.
Lu: What they’re really doing is looking at how these models process information in stages—guess, refine, and then get hit by that perturbation layer—and figuring out where the model is most confident before it gets corrupted.
Meng: From an engineering standpoint, this sounds like a clever way to keep the core logic intact even when you're relying on a very deep structure.
Lalam: It means if we can bypass that late-stage shift, we get more reliable answers in general, which could really improve how people interact with these systems daily.
Tom: Exactly. The authors show this method works across different model types, from dense structures to those Mixture-of-Experts ones they tested.
Jane: And the results are pretty strong; they’re seeing substantial reasoning gains on tough benchmarks like Omni-MATH and GPQA-Diamond using this technique.
Lu: They've even quantified the trade it makes, showing that for complex logic chains, this method aggressively cuts down on what they call a planning-pragmatics tradeoff.
Meng: I like that they showed the overhead is tiny, bounded by just scanning a few candidate layers and doing some small calculations per step. That’s important for real-world use.
Lalam: It makes sense because if the method is fast and keeps the logic sound, it’s a much better way to build culture around these tools.
Tom: So what this paper really suggests is that just making things deeper doesn't guarantee better performance when alignment is involved; you have to be smart about where you look inside the model.
Jane: It means we can stop treating those final layers as a fixed endpoint and start treating them like a potential source of error that we need to actively monitor.
Lu: The authors conclude by suggesting that using this deterministic selection method, setting it to one point, outperforms the random choices they tested.
Meng: That's a solid conclusion because it gives us a clear configuration to test right away instead of having to experiment with different settings.
Lalam: It’s exciting because if we can make reasoning more robust and reliable, the cultural impact on things like education or complex problem-solving is going to be huge.
Tom: It really does, and this opens up the next big question about how we design these models—should we just keep adding layers, or should we focus on how those layers interact?
Xuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang, Sean Du, Chujie Zheng, Fei Huang, Dayiheng Liu
Qwen Team, Alibaba Inc. · Tsinghua University
cs.CL
Submitted: 2026-06-20
Updated: 2026-10-04
Comments: EMNLP 2026 Oral
Code: https://github.com/QwenLM/Confident-Decoding
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 93/100
The gist: The paper introduces Confident Decoding to address the "Alignment Tax," a phenomenon where final-layer perturbations in large language models disrupt reasoning-relevant predictions, potentially
Key concepts
- Alignment Tax
- This is the problem where small changes made late in training disrupt a model's ability to perform complex reasoning. Instead of producing deep logic, the model defaults to generic outputs because the final layers introduce noise that derails established reasoning chains.
- Guess–Refine–Perturb Dynamic
- LLM processing follows a three-phase pattern: early layers make coarse guesses, intermediate layers refine the meaning for reasoning, and final layers cause a sharp shift. Confident Decoding targets this by stopping before the final layer's disruptive perturbation occurs.
- Entropy Guided Conservative Backward Search
- This is the core technique used to select the best layer. The method scans backward from a near-final window, looking for the first point where entropy (uncertainty) reaches a trough. This identifies a high-confidence state before potential late-layer biases take over.
- Relative Contribution Norm
- This metric is used to characterize how much each layer contributes to the model's output during processing. It helps define the three phases of LLM behavior: guessing, refining, and finally perturbing, allowing the decoder to choose a stable point.
Terminology
Summary
The paper introduces Confident Decoding to address the Alignment Tax,
a phenomenon where final-layer perturbations in large language models disrupt reasoning-relevant predictions, potentially leading to generic outputs at the expense of complex logic. This strategy employs a training-free, dynamic layer selection mechanism guided by entropy to bypass these late-stage biases while preserving semantic fidelity.
How it works
The paper identifies a recurring Guess–Refine–Perturb
dynamic in LLM forward passes: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers introduce a sharp representational shift that can move the prediction away from the most reasoning-relevant intermediate state. This suggests that standard autoregressive decoding from the final layer is not always optimal for complex tasks.
Confident Decoding is a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy guided conservative backward search
. It does not truncate the transformer or modify the model’s forward pass. Instead, it computes candidate logits from a small near-final layer window and selects the first local entropy trough encountered while scanning backward. This rule identifies the point at which the model reaches high confidence before a potential late-layer perturbation emerges.
Theoretical Grounding and Dynamics
The paper formalizes layer dynamics using two key metrics: Relative Contribution Norm
and Residual I/O Cosine Similarity
. These metrics characterize the three-phase progression of LLM forward passes: Phase I (Guess) where the first decoder layer’s contribution vector is high, Phase II (Refine) where sublayer contribution is smaller than the existing residual and directional fidelity remains high, and Phase III (Perturbation) where Norm Ratio re-elevates and IO-CosSim drops sharply.
The paper models alignment as a Tax vs. Guardrail
by defining a risk function R(l) that includes the KL divergence term DKL(Plogic∥Palign). It shows that for complex reasoning tasks, this perturbation acts as an unnatural 'entropy oscillation' that derails the established logic chain,
whereas for safety tasks, it serves as a guardrail.
Empirical Validation and Results
Experiments across dense and Mixture-of-Experts (MoE) architectures demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE. Table 1 shows that Confident Decoding consistently secures substantial reasoning gains
across all evaluated models. The Planning–Pragmatics Tradeoff
is most destructive on fragile, low-frequency logic chains, and Confident Decoding aggressively neutralizes this threat.
Computational Overhead and Robustness
Confident Decoding maintains negligible computational overhead. The extra per-step cost is bounded by one batched unembedding of M candidate hidden states (O(MBdV)) and a K-step trough scan (O(KB)), which is generally not greater than the regular basic decoding cost. The overall end-to-end wall-clock latency increase is strictly less than 2% per token. The method is shown to be architecture-agnostic, yielding positive average gains on every backbone, confirming that the Alignment Tax
perturbation is an emergent property of post-training paradigms.
Conclusion
Confident Decoding successfully isolates generation from corrupted terminal distributions by dynamically locating the monotonic entropy valley during a backward scan. This strategy proves that bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs. The deterministic valley selection setting (p=1.0) is adopted as the default configuration because it outperforms stochastic mixtures by bypassing the alignment-tax bias. Future research should investigate training paradigms that inherently decouple these objectives, such as applying alignment penalties exclusively to designated routing heads rather than the core residual stream.
REFERENCES
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarmaSarma et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021
Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong et al. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3639–3664, 2025
Nora Belrose, Igor Ostrovsky, Lev McKinney, Zach Furman, Logan Smith, Danny Halawi, Stella Biderman and Jacob Steinhardt. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023
Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou et al. Do NOT think that much for 2+3=? on the overthinking of o1-like LLMs. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 9487–9499 PMLR, 2025
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R Glass, and Pengcheng He. Dola: Decoding by contrasting layers improves factuality in large language models. In The Twelfth International Conference on Learning Representations, 2023
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021
Róbert Csordás, Christopher D Manning, and Christopher Potts. Do language models use their depth efficiently? In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
Souvik Das, Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, and Dong Yu. Entropy guided extrapolative decoding to improve factuality in large language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 6589–6600, 2025
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. In International Conference on Learning Representations, 2019
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1):12, 2021
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer et al. Layerskip: Enabling early exit inference and self-speculative decoding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12622–12642, 2024
Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng et al. Not all layers of LLMs are necessary during inference.
Improvements for AI systems
-
Confidence-based Layer Selection for Reasoning: Implement Confident Decoding to dynamically select
the most reliable near-final layer through entropy-guided conservative backward search,
which allows the model todynamically bypass final-layer perturbations
andrecover domain-specific, semantically precise terminology
when standard decoding fails. -
Alignment Tax Mitigation: Reduce the impact of post-training alignment objectives by filtering late-stage noise; this improves performance on complex reasoning benchmarks like GPQA-Diamond and Omni-MATH where the
Planning–pragmatics tradeoff
is most destructive, effectivelyneutralizing the 'Alignment Tax'.
-
Architectural Robustness Across Model Types: Ensure generalized performance across dense and Mixture-of-Experts (MoE) architectures by leveraging the finding that
The entropy valley serves as a universal, mathematical anchor for recovering semantic fidelity,
confirming gains on models like Qwen3.5-A3B and gpt-oss series. -
Optimized Computational Efficiency: Maintain production viability by employing a
surgical intervention
strategy where the overhead is bounded to "< 2% per tokenlatency increase, achieved through a
highly sparse lazy evaluation dynamic" that only incurs computation when necessary for conflict resolution.
Abstract
Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred tokens. We introduce Confident Decoding, a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy-guided conservative backward search. We further provide a theoretical formulation of layer selection as an optimal stopping problem, showing that under bounded projection noise and dominant late-stage alignment perturbation, our search rule filters perturbation while bounding the loss relative to the oracle refinement layer. Experiments across dense and Mixture-of-Experts LLMs demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE, with zero memory overhead and less than 2% latency increase. These results suggest dynamically bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs.
Sources
- A General Language Assistant as a Laboratory for Alignment
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Training Verifiers to Solve Math Word Problems
- How Do LLMs Use Their Depth?
- Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models
- The Diminishing Returns of Early-Exit Decoding in Modern LLMs
- Dynamic Early Exit in Reasoning Models
- LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering