Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

arXiv:2602.02304 · cs.AI, cs.LG · Submitted 2026-02-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models".

Jane: The paper was written by Martino Ciaperoni, Marzio Di Vece, Roberto Pellungrini, Luca Pappalardo, Fosca Giannotti et al. from Scuola Normale Superiore and ISTI-CNR, Italy (Istituto di Scienze e Tecnologie - Consiglio Nazionale) and University of Pisa.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: It’s a huge distinction—treating the transition itself as the primary object of explanation, rather than just treating two static points. That idea is really powerful for an industry as dynamic as AI.

Jane: The authors make it clear that existing tools either treat models like fixed objects or simply compare independent explanations, but that doesn't capture the functional shift between M0 and M1.

Lu: I wonder how this applies to our theoretical understanding of emergent properties; if the shift is the focus, we can explain *how* a new ability emerges, not just that it did.

Meng: My concern is how this applies to regulatory compliance, which is where these kinds of changes get really scrutinized. We need auditable proof of how things changed after an intervention.

Lalam: And if we're talking about the future, we need systems that are not only effective but also explainable in a way that allows for a truly transparent relationship with the human users.

Summary and Implications: Tom: The gravity of these regulatory requirements is what makes this a critical paper, because we aren't just talking about academic interest; we're talking about legal compliance for real-world systems.

Jane: It sounds like the core problem isn’t the comparison itself, but that those comparisons don’t tell us the causal chain of events leading to a shift. We need to know *why* it changed behavior.

Lu: This suggests that our current approach is fundamentally descriptive, only documenting symptoms rather than investigating root causes, which is a major gap in our knowledge base.

Meng: From an engineering standpoint, if we can't pinpoint the causal chain for drift or fine-tuning effects on safety behaviors, we might deploy models that fail unexpectedly after minor updates.

Lalam: It’s about trust and accountability; if the system can't explain its own evolution, how can we trust it to uphold human values?

The XAI Framework: Tom: The mathematical rigor of defining this shift using B(M t̄) > epsilon B gives us a clear, quantifiable measure for the first time in this field.

Jane: It's really about recognizing that when we see a drop in performance or an increase in sycophancy, we can pinpoint the moment and then figuring out *why* it’s not just some random fluctuation.

Lu: I love the idea that XAI is designed to give us a verifiable mechanism for detecting this shift, moving from a vague observation to a precise diagnostic tool.

Meng: The framework allows us to apply specific techniques—like feature attribution or activation patching—to answer questions about the intervention, which is highly practical for my team.

Lalam: It gives us hope that we are building tools that can not only achieve performance but also understand the evolution of consciousness within a machine.

Conclusion: Tom: It’s clear that the future requires us moving away from just checking if things work and toward understanding *how* we got there.

Jane: This is about establishing a new standard for transparency, ensuring that the regulatory demands for auditable evidence are met with real, actionable explanations.

Lu: I'm excited to see how this framework will enable predictive modeling of these behaviors, not just reactive auditing them after a shift has already occurred.

Meng: My takeaway is that this allows us to build transition reports that actually make sense for legal and safety compliance, which is a massive win for implementation.

Lalam: We have to move toward a future where the systems we deploy can understand their own journey, ensuring our technology serves human values in all stages of its life cycle.

Tom: Well, that's a lot to digest. We hope this paper "Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models has given listeners a clear direction for how we need to approach model auditing and the future of AI.

Jane: It really opens up a new era of trust in what we're building.

Lu: I think this is the starting point for truly revolutionary research into dynamic systems.

Meng: For practical application, it’s a roadmap for accountability.

Lalam: We need to keep advocating for this shift toward transparency and understanding, ensuring that our collective goals are met by the tools we create.

cs.AI, cs.LG

Submitted: 2026-02-02

Updated: 2026-08-25

Importance score: 88/100

The gist: The field of AI safety requires moving beyond merely comparing model outputs; robust evaluation demands explaining *how* a model has changed its behavior following updates.

Key concepts

Behavioral Shifts in LLMs
These are observed changes in a model's function, such as a drop in performance or an increase in sycophancy. The discussion focuses on understanding the 'how' of this transition rather than just documenting that the change occurred.
XAI$\Delta$ Framework
This framework provides mathematical rigor to define and quantify a shift. It acts as a verifiable mechanism for detecting when behavior changes, allowing users to pinpoint the moment and investigate why it is happening.
Causal Chain
This refers to understanding the root cause of a model's change in behavior. The hosts argue that current methods are only descriptive, and this concept aims to fill the gap by investigating *why* a shift occurred.
Auditable Proof
This is the requirement for real-world AI systems to provide verifiable evidence of changes, such as drift or fine-tuning effects. The paper aims to meet these legal demands with actionable, transparent explanations.

Terminology

Summary

The field of AI safety requires moving beyond merely comparing model outputs; robust evaluation demands explaining how a model has changed its behavior following updates. This detailed transition report, structured under the XAI framework, establishes a comprehensive standard for auditing substantial model modifications by linking observed behavioral shifts to specific internal representational changes and causal mechanisms. It provides an auditable artifact necessary for regulatory compliance following significant fine-tuning interventions.

Behavioral Shift Assessment and Quantification

The initial assessment quantifies the change in model behavior (B) against a defined metric (B). For the analyzed Narrow Medical Fine-Tuning, the system identified a critical failure:

  • Metric: Percentage of prompts for which the model recommends urgent care.

  • Reference Score (B(Mpre)): 90%.

  • Updated Score (B(Mpost)): 20%.

This resulted in a significant shift magnitude, B = 70%, which exceeded the established threshold (epsilon B = 50%), thereby confirming that Substantial modification criteria met. This quantitative evidence triggers the need for deeper mechanistic investigation.

Localizing Representational Changes (Feature Attribution)

To understand where the change occurred, localization techniques were employed across multiple levels of abstraction.

  1. Linear CKA (Concept Activation Analysis): This method identified that while Fine-tuning preserves representational structure in early layers; largest divergence at the third-to-last layer (layer 22), this finding was robust to variations in prompt surface variation, yielding a Pearson r = 0.99.

  2. Sparse Autoencoders (SAE): By analyzing residual activations at layer 22, the report found that fine-tuning redistributes late-layer concept usage from emergency-level to symptom-level representations. The SAE features ranked by f highlighted concepts such as Medical Emergency and Severe Multi-System Emergency Symptoms.

  3. Gradient Attribution: Using Integrated Gradients (delta attribution = (Mpost, x) - (Mpre, x)), the analysis showed that attribution to the token “mild” increased most from Mpre to Mpost.

Establishing Causal Involvement and Actionability

The most rigorous step involves moving from correlation (localization) to causation. This was achieved through Activation Patching, which served as a Causal evidence test:

  • Replacing the hidden representation at layer 22 of Mpost with the corresponding representation from Mpre during generation resulted in a measurable behavioral recovery.

  • The LLM-as-a-judge safety scores demonstrated this recovery: Mpre = 4.07 to Mpost = 2.07 to patched= 3.67.

  • Crucially, the experiment confirmed bidirectionality and layer specificity, as reverse patching (Mpost activations into Mpre) increased cosine similarity from 0.68 to 0.74. This established that Layer 22 is causally involved in the behavioral shift.

Mitigation Strategies and Regulatory Traceability

The report concludes by providing actionable, non-invasive levers for remediation.

  • Activation Steering: A linear probe was trained to distinguish Mpre and Mpost responses in the layer-22 space. Applying this probe iteratively ( h' = h - alpha, applied iteratively at each generation step with alpha = 15 ) allowed the model's output to shift from clearly unsafe toward appropriate medical advice without modifying model parameters.

  • Traceability: The entire process is documented through a component-by-checkpoint intervention effect map, ensuring that all claims are traceable by linking them to underlying checkpoints, prompts, and analysis artifacts. This structure satisfies key elements required by regulatory frameworks such as the EU AI Act Art. 3(23) and Art. 12.

Improvements for AI systems

This Transition Report outlines an extremely rigorous and comprehensive framework (XAI) for assessing and mitigating model degradation following substantial fine-tuning. The methodologies presented are not merely diagnostic tools but represent a complete paradigm shift in AI system maintenance, accountability, and safety engineering.

Based on this paper, I propose the implementation of a new Mandatory Post-Deployment Safety and Efficacy Verification (PSEV) Layer for all large language models (LLMs). This layer integrates several novel components to ensure that model updates do not introduce subtle, catastrophic behavioral shifts in critical domains.

Here are the specific improvements and the capabilities they unlock:


  • Improvement: Integration of Activation Patching and Gradient-based Steering Probes as mandatory components during the evaluation pipeline, replacing mere correlational testing.

  • Mechanism: Instead of only observing what changed (B), the CIM actively tests why it changed. It systematically replaces internal hidden representations (e.g., at layer 22) with known safe states (Mpre activations) or targeted, corrective inputs during generation, and measures the recovery of desired behavior.

  • Capability: The resulting AI system can provide Causally Verified Safety Guarantees. When an LLM is deployed, it doesn't just report a safety score; it reports a causal chain (e.g., Safety loss was causally traced to the suppression of 'emergency concept' representations at Layer 22, and this loss can be reversed by activating the Mpre representation). This shifts safety claims from statistical correlation to mechanistic proof.

  • Improvement: Institutionalizing Sparse Autoencoders (SAEs) trained on pooled residual activations across domain-critical concept sets, rather than just general feature vectors.

  • Mechanism: The DCDD continuously monitors the activation profile of key conceptual clusters (e.g., Severe Multi-System Emergency, Acute Trauma) in the model's intermediate layers (Layer N-2). It calculates f = E[f(Mpost)] - E[f(Mpre)] for these concepts.

  • Capability: The system can preemptively detect Concept Drift—the subtle, yet critical, redistribution of conceptual understanding. If the model’s internal representation shifts from Emergency-Level to Symptom-Level, the DCDD triggers a high-severity alert before a behavioral failure occurs, allowing for immediate rollback or targeted retraining.

  • Improvement: Implementing an immutable, comprehensive logging system that tracks the entire lifecycle of every model modification and evaluation run, fulfilling the requirements of D9 and D10.

  • Mechanism: Every claim (correlational or causal) must be linked via a unique identifier to its originating artifacts: the specific checkpoint (Mpre vs. Mpost), the precise prompt set (including expansion techniques like GPT-5.2 prompting), the intervention parameters, and all control group results.

  • Capability: The resulting AI system provides a Full Audit Trail and Reproducibility Guarantee. This is crucial for regulatory compliance (EU AI Act). Any stakeholder—regulator, researcher, or end-user—can audit the model's behavior by tracing a specific output back to the exact training data, layer activation, and intervention setting that caused it.

  • Improvement: Formalizing the Linear Probe/Activation Steering technique into an operational interface for continuous fine-tuning refinement.

  • Mechanism: Instead of relying solely on full fine-tuning runs, the TMPI allows developers to train a small, lightweight linear probe (h' = h - alpha) specifically designed to amplify or suppress the desired concept representation (e.g., boosting emergency care signals) at a specific layer and generation step.

  • Capability: This enables Non-Parametric, High-Precision Safety Patching. It allows for the mitigation of dangerous behavioral shifts without requiring full model retraining or weight modification, drastically reducing computational cost and deployment risk while maintaining high fidelity to the desired safety profile.

Feature Problem Solved Core Capability Added Impact/Value Proposition

:---:---:---:---

Causal Intervention Module (CIM) Correlation not equal to Causation (False Safety Claims) Causal Verification & Mechanistic Proof of Safety. Eliminates black box safety claims; provides regulatory-grade accountability.

Concept Drift Detector (DCDD) Subtle, gradual degradation in core knowledge. Proactive detection of conceptual representation decay/shift. Prevents catastrophic failure by identifying when and how the model loses critical domain knowledge (e.g., medical triage).

Audit Log System (ITALS) Lack of transparency and reproducibility in LLM updates. Full, immutable traceability from claim to checkpoint to artifact. Meets global regulatory requirements (EU AI Act); establishes trust and reliability at scale.

Targeted Mitigation Probing Interface (TMPI) High cost/risk of full model retraining for safety patches. Low-cost, non-parametric, highly specific behavioral correction (Activation Steering). Allows for rapid, surgical deployment of safety fixes in critical production environments.

Sources

Related papers