Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

arXiv:2610.00024 · cs.CV, cs.CL, cs.LG · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Encoded but Disconnected".

Jane: Across three vision–language model architectures, this work reports a universal negative finding for mid-layer interpretability, demonstrating that while mid-layers often encode ground-truth answers in errors,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Jane, we’ve been looking at the technical details of this paper, "Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null," and I’m really trying to nail down what it means for us right now. It sounds like they're tackling a deep issue in how we interpret these big vision models.

Jane: I totally get that, Tom; it seems like the core idea is showing that while middle layers look like they are holding the right answer, that signal isn't actually what drives the final prediction. It’s about separating what looks correct from what actually causes the error in a very specific way.

Lu: The decomposition into three distinct failure modes—Perception Failure, Encoded-but-Disconnected, and Prior-Override—is fascinating because it gives us a concrete way to categorize every mistake we see across different architectures. It moves us beyond just saying "the model got it wrong" to understanding *why* it got it wrong mechanistically.

Meng: From my side, I'm thinking about the practical engineering implications; if we can reliably classify an error this way, we can build systems that apply the right fix instantly instead of trying every possible general solution. It’s less guesswork and more targeted engineering for specific failure types.

Lalam: I think this is really interesting because it shifts our focus from just making the model perform better overall to understanding the internal mechanics of its failures, which helps us build a more resilient and predictable AI culture. If we know where the signal goes, we can design safeguards that are smarter than just patching things up blindly.

Tom: Exactly! And to add to what Lu mentioned, they found that this decomposition is quite robust; it learns above sixty percent accuracy on all three models, which shows this isn't some fluke in the test set.

Jane: That robustness is key because it confirms that these failure modes are genuinely separable and not just random correlations we accidentally stumbled upon in the data. It’s a systematic way of looking at model errors.

Lu: And what really blows me is their finding on causality; they used an activation patching protocol, and it showed zero non-trivial flips at the layer level for all three architectures, which directly challenges the idea that any information we can see in those layers is causally relevant to the final output.

Meng: That null result on patching is a big deal for us because it means we can stop wasting time trying to debug every single intermediate layer just because something looks active there; it's not driving the final decision.

Title and authors: Lalam: It’s like moving from looking at every wire in a massive machine to only checking the main power lines that actually feed the output, which makes sense for building reliable AI systems.

Tom: So, if we look at this paper, "Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null," they are showing us that the mid-layer signal is a strong predictor of failure mode, but it isn't the causal mechanism itself.

Jane: Right. The main summary points are that they define an operational way to identify three failure modes, and then use an architecture’s prior direction to predict which specific intervention will actually work for that error category under ideal conditions.

Lu: The sweet zone concept is another important piece; they defined this as a layer window where the ground-truth answer token is encoded at mid-layer but the final prediction still misses it, and this window shifts depending on the architecture, like it's about one third of the layers for LLaVA and two thirds for InternVL3.

Meng: From an engineering standpoint, knowing that this sweet zone depth varies by model means we can create a tailored scanning protocol instead of using a one-size-fits-all window size across all our different AI stacks.

Lalam: That level of specificity is what makes the work valuable for us; it gives us blueprints for how to diagnose and fix specific model types with precision, which will make our entire AI infrastructure much more predictable.

Tom: It really boils down to this: while the phenomenon of encoding the ground-truth answer in those layers is a strong correlate of failure, it isn't actually where the model generates its final prediction.

Jane: Precisely; that’s what moves us past just observing symptoms and into understanding the underlying system dynamics. It suggests we need a new way to think about interpretability based on per-sample routing rather than just population diagnostics.

Lu: And they did go further by identifying an architecture-specific causal head in InternVL3, layer twenty head two which is localized and exhibits self-attending behavior with a very low probability of less than one ten to the power of negative four.

Meng: That specific finding on that internal component is interesting for architectural design; it points to certain specialized aggregation steps that are unique to that model structure.

Lalam: If we can find these specific causal hotspots, we can develop highly targeted optimizations for those parts of the model rather than just retraining the whole thing.

Tom: And they showed that this categorization allows them to map each architecture’s dominant failure mode to a distinct intervention strategy, like prompt debiasing or mid-layer re-promotion, based on that architecture's prior direction.

Title and authors: Jane: So, it’s not just about finding the error; it’s about predicting the correct treatment for that specific error category based on where the model is leaning in its decision process.

Lu: The framework they propose allows us to route samples based on their predicted failure mode to an intervention that aligns with the architecture’s prior direction, which confirms that under oracle labels, each architecture admits a category-targeted response with a flip rate well above its prior-aligned class.

Meng: That prediction of which intervention matches the architecture seems like it’s the most actionable part for deployment; we can automate that routing decision once we have those classifiers trained.

Lalam: It means we aren't just applying generic fixes anymore; we are using the model's own internal tendencies to select the most appropriate correction, which is a step toward truly adaptive AI.

Tom: So, to wrap up this discussion on "Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null," the big implication is that we have moved from population-level diagnosis to per-sample routing based on predicted failure modes and architecture-specific prior directions.

Jane: It fundamentally changes how we approach debugging; instead of asking where the signal goes generally, we ask how this specific error mode manifests and then apply a fix tailored to that manifestation.

Lu: The paper confirms that while the phenomenology of seeing ground-truth encoding in mid-layers is universal, its causal locus isn't; it shifts depending on the backbone, and they found an architecture-specific causal head that acts differently than the others.

Meng: And for practical application, this means we can stop assuming lens accessibility equals causality and instead build systems that explicitly route around those assumed connections.

Lalam: Ultimately, this research gives us a much more precise toolkit for understanding and steering these complex vision-language models in a way that respects their internal structure rather than just treating them as black boxes.

Tom: That’s all for this paper on "Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null." We’ve seen how they decompose failures and move toward prior-aware routing.

Jane: It's been really illuminating to see how they separate the observed correlation from the actual causal mechanism in these complex systems.

Lu: I think this work opens up so many avenues for future research, especially around understanding those architecture-specific causal heads we discovered.

Meng: We should keep an eye on how this routing framework performs when we apply it to our proprietary models; that’s where the real test will be for practical impact.

Lalam: I'm really looking forward to seeing how this per-sample routing can help us make our AI more robust and reliable in the long run.

The paper's summary: Tom: So, to recap what we've discussed about "Encoded but Disconnected," the core finding is that while we can see ground truth information encoded in some middle layers of a vision language model when it makes a mistake, that signal doesn't actually cause the final output prediction.

Jane: That’s exactly right, Tom; it shows that those mid-layers are just holding onto things temporarily, like they’re trying to remember the answer from the training data, but they aren't driving the actual decision-making process at all.

Lu: I think what really hits home is that this isn't a single failure mode we’re dealing with; instead, it breaks down every error into three distinct categories—Perception Failure, Encoded-but-Disconnected, and Prior-Override—which gives us a much richer way to diagnose problems.

Meng: From an engineering standpoint, that means we can stop looking at every intermediate layer for a single mistake; we can actually classify *why* it failed first. That moves us from broad debugging to specific problem solving.

Lalam: And the most powerful part is their suggestion for a routing framework; they use the model's inherent "prior direction" to predict which specific fix will actually work for that particular error type under ideal conditions.

Tom: It’s about shifting our focus from just observing what information is visible to understanding how a specific error mode manifests internally, and then applying a tailored correction based on the model's own internal tendencies.

Jane: That's a great way to put it; it’s moving us from asking "what information is visible?" to "how does this specific error mode manifest?" This makes interpretability much more actionable for developers.

Lu: And they did show that this decomposition is quite robust, with a high accuracy rate across different architectures, confirming these failure modes are genuinely separable and not just random noise in the data.

Meng: I’m curious about the practical application of that prior-aware routing; if we can automate the decision of which intervention to use based on the architecture's direction, that could save a lot of manual tuning time.

Lalam: It suggests a future where AI systems don't just have one fix in their toolbox, but dynamically select the most appropriate correction based on the specific type of mistake they are making.

Tom: Absolutely; it’s like giving the AI a sophisticated internal decision-making layer for fixing its own errors, instead of treating every error as a generic glitch.

Jane: It really helps us build systems that are more predictable because we know exactly which repair strategy is likely to succeed given the model's current state and failure type.

Lu: I'm even excited about the architecture-specific causal head they identified in InternVL3; finding those unique self-attending components points toward specialized optimization opportunities for certain model backbones.

Meng: That’s something I want to look into; identifying those hotspots means we can focus our resource allocation on the parts of the network that are actually doing the heavy lifting, rather than guessing where to apply a general patch.

Lalam: If we can pinpoint those unique causal mechanisms, it helps us build a more nuanced understanding of how different model designs behave when they fail.

Tom: So, we've got this framework that lets us diagnose the error type and then choose the right fix based on the model’s internal direction; it’s a significant step in making AI debugging much more precise.

The paper's improvements: Tom: So, if we look at what the authors suggest for improving these systems, they aren't just looking for bigger models; they're suggesting a complete shift in how we debug and fix errors by incorporating that prior-aware routing framework into the deployment pipeline.

Jane: That makes sense; it means instead of having a general repair kit, the system should be smart enough to look at what kind of error it just made and automatically select the most suitable intervention from a menu.

Lu: I think this moves us toward truly adaptive AI because it means we're not applying a one-size-fits-all solution across all models; instead, we tailor the fix based on the specific failure mode they classify.

Meng: From my standpoint as an engineer, that automation is key; if the system can predict which intervention—like prompt debiasing or mid-layer re-promotion—is needed for a particular error category, we can build automated guardrails that reduce human intervention drastically.

Lalam: That level of adaptability in the AI's self-correction process really speaks to improving our overall AI culture; it shows we’re building systems that are not just static tools but evolving entities capable of learning how to recover from specific types of mistakes.

Tom: It’s about moving past simple error correction and toward a much more sophisticated, context-aware repair mechanism for vision language models.

Jane: And the authors emphasize that this routing prediction works best under oracle labels, which is important because it grounds the framework in reality when we have perfect ground truth to train on.

Lu: The paper suggests that as future work, researchers should focus heavily on those architecture-specific causal heads they found; understanding exactly what those unique components are doing could open up entirely new avenues for model architecture design.

Meng: Identifying those specialized causal components means we can target our optimization efforts precisely where they matter most in specific model families, which is a lot more efficient than broad retraining.

Lalam: I think focusing on those specialized mechanisms helps us understand the underlying logic of different AI structures, which ultimately makes our entire ecosystem more robust and trustworthy for users.

Tom: So, the paper's takeaway isn't just a diagnostic tool; it’s a blueprint for building next-generation AI that can self-diagnose its own errors and apply highly specific, contextually appropriate remedies.

Conclusion: Tom: So, to wrap up this whole discussion on "Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null," we’ve seen how they move beyond just spotting errors to understanding the specific mechanism causing those errors in vision language models.

Jane: Exactly; they give us a robust way to categorize failures into distinct modes, which allows us to stop treating every mistake as the same problem and start applying targeted solutions.

Lu: This work opens up so many creative possibilities for how we design next-generation AI systems because it shows that internal representations have much more nuanced behavior than we previously thought when things go wrong.

Meng: I think the real impact here is in building more reliable deployment pipelines; if we can route samples based on these failure modes, it means our operational guardrails can become significantly smarter and less prone to false positives.

Lalam: For me, this research is really about improving the culture of AI development by showing us that deep introspection into failure modes leads to a more resilient and thoughtful way of building and trusting these complex systems.

Tom: It’s been fascinating seeing how they link that specific error mode to an architecture-specific intervention strategy, which is a huge step toward making our AI deployments truly adaptive.

Jane: It really reinforces the idea that interpretability isn't just about looking at the surface; it’s about understanding the dynamic process happening inside when something goes wrong.

Lu: And their findings on those architecture-specific causal heads suggest we need to start thinking more structurally about how different model backbones handle these kinds of internal signals.

Meng: I agree, focusing on those specialized components helps us optimize resources better and prevents us from wasting time debugging parts of the network that aren't actually driving the final prediction.

Lalam: Knowing exactly how different models fail allows us to build systems that are not just accurate but also much more reliable and trustworthy in real-world applications.

Tom: So, we’ve seen how this paper on "Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null" gives us a clearer path from observing symptoms to designing precise, prior-aware countermeasures.

Jane: It's been really illuminating seeing how they separate the observed correlation from the actual causal mechanism in these complex systems.

Lu: I think this work opens up so many avenues for future research, especially around understanding those architecture-specific causal heads we discovered and how they might relate to other model structures.

Meng: We should keep an eye on how this routing framework performs when we apply it to our proprietary models; that’s where the real test will be for practical impact.

Lalam: I'm really looking forward to seeing how this per-sample routing can help us make our AI more robust and reliable in the long run.

Genpei Zhang

University of Wisconsin–Madison

cs.CV, cs.CL, cs.LG

Submitted: 2026-09-01

Updated: 2026-09-01

Comments: 13 pages, 4 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Across three vision–language model architectures, this work reports a universal negative finding for mid-layer interpretability, demonstrating that while mid-layers often encode ground-truth

Key concepts

Encoded Answer Signal (EAS)
This phenomenon occurs when a model successfully encodes the correct ground-truth answer token at a specific layer within its architecture's 'sweet zone.' However, despite this internal encoding, the final prediction made by the model does not match that encoded answer. It serves as a strong indicator of an error state.
Activation Patching Protocol
This experimental method tests causality by replacing a residual stream in the model with data from a 'donor' sample where the model is already correct. The results showed no significant change in prediction flips at the layer level across all tested architectures, challenging the idea that information accessible via these patches is causally relevant for the final prediction.
Prior-Direction-Aware Routing
This framework uses a classifier trained on features from a single forward pass to identify which of three failure modes (Perception Failure, Encoded-but-Disconnected, or Prior-Override) occurred. Based on this prediction, it suggests a specific intervention—like prompt debiasing or mid-layer re-promotion—that matches the model's expected prior direction for that category.
Causal Head
This refers to a specific, localized mechanism within a model architecture that exhibits non-vocabulary, self-attending behavior. In one architecture tested, this was found at a precise layer and head with very low probability. This suggests the signal is not direct answer encoding but rather an aggregated residual stream effect.

Terminology

Summary

Across three vision–language model architectures, this work reports a universal negative finding for mid-layer interpretability, demonstrating that while mid-layers often encode ground-truth answers in errors, this signal is not causally active for the final prediction. This research decomposes VLM failures into mechanistically distinct modes and proposes a prior-direction-aware routing framework that predicts which specific intervention elicits a category-specific response.

How it works

The paper introduces three failure modes—Perception Failure, Encoded-but-Disconnected, and Prior-Override—which partition every error based on two binary observations: whether the Encoded Answer Signal (EAS) condition is met in mid-layers and whether the emitted token aligns with the model’s training prior. This decomposition is operational, defined by observable mid-layer and final-layer behavior on a single forward pass. The EAS phenomenon occurs when a model encodes the ground-truth answer token at some layer within an architecture-specific “sweet zone,” but the final prediction does not match it.

How it works

The core experimental validation involves an Activation Patching Protocol to test causality. This protocol applies standard cross-instance activation patching by replacing a residual stream with that from a donor sample on which the model is correct, and re-running from the patched layer onward. The finding is that residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level. This challenges the assumption that lens-accessible information is causally relevant for final prediction.

How it works

The study identifies an architecture-specific causal head: InternVL3 layer 20, head 2, which exhibits a nonvocab, self-attending behavior localized to that specific head with a probability of "p < 10−4." This mechanism is not a direct yes/no vocab encoder; it projects the raw logit lens output through the final layernorm and unembedding matrix, attending almost entirely to its own position. This suggests the signal is aggregated residual-stream information rather than encoding the answer directly.

How it works

A per-sample classifier is trained on a 24-dimensional feature vector extracted from a single forward pass, which combines per-layer top-1/2/3 logit values and margins across six representative sweet-zone layers. This classifier separates the three failure categories with high accuracy (ranging from 68.0% to 82.5% under logistic regression). Crucially, this classifier is shown not to be trivially reconstructible by only using the two label-defining bits (EAS-hit and prior-direction), confirming that the continuous features carry non-trivial signal.

How it works

The paper demonstrates a prior-directionaware mitigation framework where each architecture’s dominant failure mode admits a different per-category intervention. For instance, prompt-debias is applied to the Prior-Override class (PO) on under-affirmative models (Qwen, InternVL3), while mid-layer RePromote is used for Encoded-but-Disconnected (EBD) errors on the over-affirmative LLaVA model. This allows the system to route samples based on their predicted failure mode to an intervention that matches the architecture’s prior direction, yielding a category-specific response.

How it works

The results confirm that under oracle category labels, each architecture admits a category-targeted intervention with a flip rate well above its prior-aligned class, and this intervention is predicted by the architecture’s prior direction. The overall conclusion is that while the mid-layer top-1 encoding of the ground-truth answer in the EBD category is a robust phenomenological correlate of failure, it is not the locus of the model’s prediction-generation mechanism. This moves interpretability from population-level diagnosis to per-sample, prior-aware routing.

The gist

The mid-layer top-1 encoding of the ground-truth answer in the EBD category, while a robust phenomenological correlate of failure, is not the locus of the model’s prediction-generation mechanism. This is a strong negative finding and directly tests the implicit assumption that lens-accessible information is causally relevant for final prediction. The decomposition is still useful: separable from a 24-feature classifier, it splits POPE failures into two mechanistically distinct halves and predicts which of two interventions matches a given architecture. The phenomenology is universal but its locus is not—the sweet-zone depth shifts with backbone (∼1/3, ∼1/2, ∼2/3), and the per-head causal effect on InternVL3 has no analog on the other two.

Summary of Findings

Improvements for AI systems

Based on this scientific paper, here are the specific improvements for AI systems and what those improved systems can achieve:


  1. Improvements based on a Single-Forward-Pass Decomposition (PF, EBD, PO) and Per-Sample Classification:

The system should be augmented with a diagnostic pipeline that performs a single forward pass to classify every error into one of three mechanistic failure modes: Perception Failure (PF), Encoded-but-Disconnected (EBD), or Prior-Override (PO).

The improved AI system can achieve:

2.1. Diagnostic Routing: The system can route errors to the correct mitigation strategy based on the classification, effectively moving from population-level diagnosis to per-sample, prior-aware routing.

2.2. Targeted Intervention: Instead of applying a uniform fix, the system applies a category-specific intervention (e.g., prompt debiasing for PO errors vs. mid-layer re-promotion for EBD errors). This allows the system to address the specific root cause of its failure, leading to more targeted and potentially more effective corrections than general decoding fixes.

  1. Improvements based on Causal Null Findings (Activation Patching Protocol):

The system's interpretation of intermediate representations must be refined by incorporating a causal null finding: mid-layer encoding of the ground-truth answer is not causally active for the final prediction under standard patching protocols.

The improved AI system can achieve:

2.3. Refined Interpretability Assumptions: The system will no longer assume that any observed mid-layer signal is causally relevant to the final output simply because it correlates with the answer. This prevents lens-accessible information fallacy in mechanistic analysis, leading to more rigorous and less overconfident claims about causal roles in intermediate layers.

  1. Improvements based on Architecture-Specific Causal Heads (InternVL3 L20 H2):

The system's architecture can be designed to recognize and isolate architecture-specific causal pathways, such as the InternVL3 layer 20 head 2, which is non-vocabulary and self-attending.

The improved AI system can achieve:

2.4. Specialized Mechanism Detection: The system can identify specific hotspots or causal heads that are not standard vocabulary encoders but aggregate residual stream information in specific positions (like the last-text position), allowing for architecture-specific debugging and optimization that targets these specialized aggregation steps.

  1. Improvements based on Prior-Direction Aware Mitigation Framework:

The system's decoding strategy should be governed by a classifier that predicts which of two prior-direction interventions (prompt debiasing or mid-layer re-promotion) is appropriate for the failure mode identified in step 2.1.

The improved AI system can achieve:

2.5. Category-Conditional Decoding: The system can dynamically switch its intervention mechanism based on the predicted failure category, ensuring that the chosen fix (e.g., input modification vs. logit blending) is mechanistically appropriate for that failure type, maximizing the chance of success under oracle conditions.

In summary, this paper enables a shift from what information is visible? to how does this specific error mode manifest? and allows for the application of highly specialized fixes tailored to the underlying architectural failure mechanism.

Abstract

Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p < 1e-4). Despite the null, the errors separate operationally into three failure modes -- Perception Failure, Encoded-but-Disconnected, Prior-Override -- learnable above 60% on all three architectures, and the architecture's prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.

Sources

Related papers