Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

arXiv:2606.21906 · cs.CL · Submitted 2026-06-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Deeper is Not Always Better".

Tom: The paper introduces Confident Decoding to address the "Alignment Tax," a phenomenon where final-layer perturbations in large language models disrupt reasoning-relevant predictions,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we're looking at this paper called "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding." It tackles this idea that just making the model deeper always makes it smarter, and claims that it doesn't actually hold up when you consider how AI models are trained.

Jane: Right, so they're pointing out a problem where those final layers of a large language model can actually mess things up. They suggest that these later layers introduce some kind of noise that pushes the model away from the actual reasoning it should be doing, sometimes toward generic answers instead of complex logic.

Lu: What's really interesting here is their analysis of what they call a "Guess–Refine–Perturb" pattern in how these models process information during a forward pass. It seems like early layers make some initial guesses, middle layers fine-tune the reasoning part, and then the final layers introduce that shift.

Meng: So if this perturbation happens at the end, it means we can't just rely on whatever that last layer spits out for our most important answers anymore because it might be a weak signal.

Lalam: That makes sense when you think about how I work; I try to keep my core reasoning solid, but if the final output layer gets too noisy, it can make my complex logic seem flimsy.

Tom: Exactly, and that’s where this paper comes in with Confident Decoding. It proposes a training-free way to bypass those late-stage biases by dynamically picking a reliable layer rather than just going straight to the end.

Jane: They describe this strategy as using an entropy guided conservative backward search to pick the best near-final layer, which means it's not retraining anything or changing how we run the model at all.

Lu: It’s quite clever because they’re not truncating the transformer or messing with the forward pass structure; they are just smart about where to look within it based on certain dynamics.

Paper summary: Meng: So, how does this dynamic selection actually work in practice when you have these complex layer contributions and residual similarities? We need to know if this is just theoretical fluff.

Tom: Well, they use metrics like Relative Contribution Norm and Residual I/O Cosine Similarity to formally characterize those three phases of the process—Guess, Refine, and Perturb—as detailed in "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding."

Jane: The paper lays out a risk function, R(l), which shows that this late-stage perturbation acts differently depending on what the AI is trying to do.

Lu: They show that for safety tasks, like following simple guardrails, this noise might actually be helpful; but for complex reasoning, it seems to derail the established logic chain instead.

Tom: That distinction between "Tax vs. Guardrail" is a really important conceptual move here. It suggests that what we consider an alignment problem for one kind of task might not be the same for another.

Meng: So if we're focused on complex reasoning, this paper claims Confident Decoding secures substantial gains on difficult benchmarks like GPQA-Diamond and Omni-MATH across different architectures.

Jane: They validate this by showing that the method consistently secures substantial reasoning gains across all models they tested, which is a strong empirical result in "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding."

Lu: The authors specifically point out that this strategy aggressively neutralizes what they call the "Planning–Pragmatics Tradeoff," which they say is most destructive when dealing with fragile, low-frequency logic chains.

Tom: That sounds like it directly addresses one of the main weaknesses we see in current LLMs when they try to handle multi-step reasoning problems.

Jane: And computationally, they’re pretty good about that; the overhead is small because you’re just doing a few extra calculations, bounded by one batched unembedding of M candidate hidden states and a K-step scan.

Meng: So if the overhead is less than two percent per token latency increase, that makes it very practical for deployment in real applications without slowing everything down significantly.

Paper summary: Lalam: From my side, having a deterministic way to select the most reliable layer means I can trust the output more consistently, which is a big plus for cultural applications where reliability matters a lot.

Tom: It really comes down to this: Confident Decoding successfully isolates the generation process from those corrupted terminal distributions by finding that monotonic entropy valley during that backward scan.

Jane: The authors conclude that this dynamic selection proves we can unlock stronger reasoning behavior from models that are already aligned, even when there’s an alignment tax present in the final layers of their structure.

Lu: They default to a deterministic setting for this valley selection because it outperforms stochastic mixtures by bypassing that alignment-tax bias, suggesting a very direct path forward.

Meng: For practical implementation right now, this means we can start testing models on those challenging benchmarks and expect better performance without having to overhaul the entire model architecture or retraining it from scratch.

Tom: So to wrap up this discussion on "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding," we see a method that isolates generation from those final layer problems by smartly navigating the internal dynamics of how an AI generates text.

Jane: It’s about finding that point where the model has high confidence before any potential late-layer perturbation kicks in, which is a really direct way to improve complex reasoning capabilities.

Lu: Future work they suggest could look at training paradigms where you apply alignment penalties only to specific routing heads instead of the entire core residual stream, which would be a different approach entirely.

Meng: That sounds like a more structural change to how we design the AI itself rather than just fixing the decoding step.

Tom: Yeah, that's interesting because it moves the problem from runtime correction into model training design, but for now Confident Decoding gives us a solid tool to boost performance on existing models.

Jane: So this paper shows that depth isn't automatically better when alignment is involved; we have to be smarter about where we look in the model's layers.

Conclusion: Tom: So we've been looking at how these models can get stuck in traps where their final layers mess up the actual reasoning, and this paper is called "Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding."

Jane: Yeah, it’s basically about finding a way to bypass that noise at the end of the generation process without having to retrain everything.

Lu: What they’re really doing is looking at how these models process information in stages—guess, refine, and then get hit by that perturbation layer—and figuring out where the model is most confident before it gets corrupted.

Meng: From an engineering standpoint, this sounds like a clever way to keep the core logic intact even when you're relying on a very deep structure.

Lalam: It means if we can bypass that late-stage shift, we get more reliable answers in general, which could really improve how people interact with these systems daily.

Tom: Exactly. The authors show this method works across different model types, from dense structures to those Mixture-of-Experts ones they tested.

Jane: And the results are pretty strong; they’re seeing substantial reasoning gains on tough benchmarks like Omni-MATH and GPQA-Diamond using this technique.

Lu: They've even quantified the trade it makes, showing that for complex logic chains, this method aggressively cuts down on what they call a planning-pragmatics tradeoff.

Meng: I like that they showed the overhead is tiny, bounded by just scanning a few candidate layers and doing some small calculations per step. That’s important for real-world use.

Lalam: It makes sense because if the method is fast and keeps the logic sound, it’s a much better way to build culture around these tools.

Tom: So what this paper really suggests is that just making things deeper doesn't guarantee better performance when alignment is involved; you have to be smart about where you look inside the model.

Jane: It means we can stop treating those final layers as a fixed endpoint and start treating them like a potential source of error that we need to actively monitor.

Lu: The authors conclude by suggesting that using this deterministic selection method, setting it to one point, outperforms the random choices they tested.

Meng: That's a solid conclusion because it gives us a clear configuration to test right away instead of having to experiment with different settings.

Lalam: It’s exciting because if we can make reasoning more robust and reliable, the cultural impact on things like education or complex problem-solving is going to be huge.

Tom: It really does, and this opens up the next big question about how we design these models—should we just keep adding layers, or should we focus on how those layers interact?

Xuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang, Sean Du, Chujie Zheng, Fei Huang, Dayiheng Liu

Qwen Team, Alibaba Inc. · Tsinghua University

cs.CL

Submitted: 2026-06-20

Updated: 2026-10-04

Comments: EMNLP 2026 Oral

Code: https://github.com/QwenLM/Confident-Decoding

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 93/100

The gist: The paper introduces Confident Decoding to address the "Alignment Tax," a phenomenon where final-layer perturbations in large language models disrupt reasoning-relevant predictions, potentially

Key concepts

Alignment Tax
This is the problem where small changes made late in training disrupt a model's ability to perform complex reasoning. Instead of producing deep logic, the model defaults to generic outputs because the final layers introduce noise that derails established reasoning chains.
Guess–Refine–Perturb Dynamic
LLM processing follows a three-phase pattern: early layers make coarse guesses, intermediate layers refine the meaning for reasoning, and final layers cause a sharp shift. Confident Decoding targets this by stopping before the final layer's disruptive perturbation occurs.
Entropy Guided Conservative Backward Search
This is the core technique used to select the best layer. The method scans backward from a near-final window, looking for the first point where entropy (uncertainty) reaches a trough. This identifies a high-confidence state before potential late-layer biases take over.
Relative Contribution Norm
This metric is used to characterize how much each layer contributes to the model's output during processing. It helps define the three phases of LLM behavior: guessing, refining, and finally perturbing, allowing the decoder to choose a stable point.

Terminology

Summary

The paper introduces Confident Decoding to address the Alignment Tax, a phenomenon where final-layer perturbations in large language models disrupt reasoning-relevant predictions, potentially leading to generic outputs at the expense of complex logic. This strategy employs a training-free, dynamic layer selection mechanism guided by entropy to bypass these late-stage biases while preserving semantic fidelity.

How it works

The paper identifies a recurring Guess–Refine–Perturb dynamic in LLM forward passes: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers introduce a sharp representational shift that can move the prediction away from the most reasoning-relevant intermediate state. This suggests that standard autoregressive decoding from the final layer is not always optimal for complex tasks.

Confident Decoding is a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy guided conservative backward search. It does not truncate the transformer or modify the model’s forward pass. Instead, it computes candidate logits from a small near-final layer window and selects the first local entropy trough encountered while scanning backward. This rule identifies the point at which the model reaches high confidence before a potential late-layer perturbation emerges.

Theoretical Grounding and Dynamics

The paper formalizes layer dynamics using two key metrics: Relative Contribution Norm and Residual I/O Cosine Similarity. These metrics characterize the three-phase progression of LLM forward passes: Phase I (Guess) where the first decoder layer’s contribution vector is high, Phase II (Refine) where sublayer contribution is smaller than the existing residual and directional fidelity remains high, and Phase III (Perturbation) where Norm Ratio re-elevates and IO-CosSim drops sharply.

The paper models alignment as a Tax vs. Guardrail by defining a risk function R(l) that includes the KL divergence term DKL(Plogic∥Palign). It shows that for complex reasoning tasks, this perturbation acts as an unnatural 'entropy oscillation' that derails the established logic chain, whereas for safety tasks, it serves as a guardrail.

Empirical Validation and Results

Experiments across dense and Mixture-of-Experts (MoE) architectures demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE. Table 1 shows that Confident Decoding consistently secures substantial reasoning gains across all evaluated models. The Planning–Pragmatics Tradeoff is most destructive on fragile, low-frequency logic chains, and Confident Decoding aggressively neutralizes this threat.

Computational Overhead and Robustness

Confident Decoding maintains negligible computational overhead. The extra per-step cost is bounded by one batched unembedding of M candidate hidden states (O(MBdV)) and a K-step trough scan (O(KB)), which is generally not greater than the regular basic decoding cost. The overall end-to-end wall-clock latency increase is strictly less than 2% per token. The method is shown to be architecture-agnostic, yielding positive average gains on every backbone, confirming that the Alignment Tax perturbation is an emergent property of post-training paradigms.

Conclusion

Confident Decoding successfully isolates generation from corrupted terminal distributions by dynamically locating the monotonic entropy valley during a backward scan. This strategy proves that bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs. The deterministic valley selection setting (p=1.0) is adopted as the default configuration because it outperforms stochastic mixtures by bypassing the alignment-tax bias. Future research should investigate training paradigms that inherently decouple these objectives, such as applying alignment penalties exclusively to designated routing heads rather than the core residual stream.

REFERENCES

Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarmaSarma et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021

Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong et al. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3639–3664, 2025

Nora Belrose, Igor Ostrovsky, Lev McKinney, Zach Furman, Logan Smith, Danny Halawi, Stella Biderman and Jacob Steinhardt. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023

Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou et al. Do NOT think that much for 2+3=? on the overthinking of o1-like LLMs. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 9487–9499 PMLR, 2025

Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R Glass, and Pengcheng He. Dola: Decoding by contrasting layers improves factuality in large language models. In The Twelfth International Conference on Learning Representations, 2023

Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021

Róbert Csordás, Christopher D Manning, and Christopher Potts. Do language models use their depth efficiently? In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

Souvik Das, Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, and Dong Yu. Entropy guided extrapolative decoding to improve factuality in large language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 6589–6600, 2025

Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. In International Conference on Learning Representations, 2019

Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1):12, 2021

Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer et al. Layerskip: Enabling early exit inference and self-speculative decoding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12622–12642, 2024

Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng et al. Not all layers of LLMs are necessary during inference.

Improvements for AI systems

  1. Confidence-based Layer Selection for Reasoning: Implement Confident Decoding to dynamically select the most reliable near-final layer through entropy-guided conservative backward search, which allows the model to dynamically bypass final-layer perturbations and recover domain-specific, semantically precise terminology when standard decoding fails.

  2. Alignment Tax Mitigation: Reduce the impact of post-training alignment objectives by filtering late-stage noise; this improves performance on complex reasoning benchmarks like GPQA-Diamond and Omni-MATH where the Planning–pragmatics tradeoff is most destructive, effectively neutralizing the 'Alignment Tax'.

  3. Architectural Robustness Across Model Types: Ensure generalized performance across dense and Mixture-of-Experts (MoE) architectures by leveraging the finding that The entropy valley serves as a universal, mathematical anchor for recovering semantic fidelity, confirming gains on models like Qwen3.5-A3B and gpt-oss series.

  4. Optimized Computational Efficiency: Maintain production viability by employing a surgical intervention strategy where the overhead is bounded to "< 2% per token latency increase, achieved through a highly sparse lazy evaluation dynamic" that only incurs computation when necessary for conflict resolution.

Abstract

Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred tokens. We introduce Confident Decoding, a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy-guided conservative backward search. We further provide a theoretical formulation of layer selection as an optimal stopping problem, showing that under bounded projection noise and dominant late-stage alignment perturbation, our search rule filters perturbation while bounding the loss relative to the oracle refinement layer. Experiments across dense and Mixture-of-Experts LLMs demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE, with zero memory overhead and less than 2% latency increase. These results suggest dynamically bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs.

Sources

Related papers