Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Encoded but Not Routed".
Tom: Multimodal Large Language Models (LLMs) are increasingly used for scientific peer review, yet they exhibit a significant performance gap when verifying claims supported by charts compared to tables,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, what we just discussed was the setup for this study investigating "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." The authors are Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, and Florian Boudin. This paper is really zeroing in on a specific performance disparity where models perform better on tables than charts for scientific claim verification.
Jane: It’s about trying to figure out if the problem is that the AI can't read the chart information at all, or if it can read it but doesn't know how to use that visual data when making its final judgment. The title suggests they are investigating this routing aspect, which is a really interesting angle.
Lu: Exactly! They are testing whether this gap comes from a failure in extracting the visual information itself, or if it stems from a failure in routing that extracted information to the prediction stage, which is what they hypothesize. It’s shifting the focus from perception to utilization.
Meng: A routing failure sounds like a bottleneck in the model's internal architecture, perhaps how different types of evidence—like tables versus charts—are prioritized or passed along through those layers. That makes me think about how we structure our attention mechanisms in these large models.
Lalam: I agree with Lu; if the signal is recoverable from intermediate representations but doesn't reach the prediction position, that's a clear indication that the mechanism responsible for connecting visual evidence to the final decision needs adjustment. It’s about optimizing the flow, not just increasing the raw capacity of each component.
The paper's summary: Tom: Moving on, let's talk about what they actually found in "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." The core finding is quite striking: chart-relevant signal is indeed encoded within the models’ intermediate representations, but it fails to reach the prediction position. This suggests a failure in routing rather than a deficiency in initial encoding of that visual data.
Jane: So, to put it simply, the paper discovered that when you look deep inside the AI's processing layers, you can find evidence of what's happening in a chart, but by the time the model gets to making its final guess about whether a claim is true or false, that specific visual signal is missing from its decision path.
Lu: That confirms their hypothesis regarding routing failure. They used linear probing and attention analysis to show this pattern across three different open-weight multimodal LLMs on the SciTabAlign+ benchmark, which pairs claims with both tables and various chart types. The results were consistent, showing this same pattern regardless of the specific model family they tested.
Meng: That consistency is important for us; it means this isn't just an anomaly in one specific model; it points to a fundamental challenge in how current multimodal AI systems handle visual evidence compared to tabular data structures. It suggests we need a more universal way to integrate different modalities effectively.
Lalam: From my perspective, this summary confirms that the problem isn't that the models are blind to charts; they are blind to routing them correctly into the final decision-making process, which is a much more tractable problem for refinement than trying to teach them how to see better.
The paper's improvements: Tom: Now, let’s discuss what the authors suggest as improvements based on their findings in "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." They are not just pointing out a problem; they are suggesting that we need to look at how models utilize these internal signals more effectively.
Jane: It seems like the paper points toward diagnosing this issue through specific probing techniques, like looking at whether information aggregates sharply at the last token position for tables versus remaining scattered across layers for charts. This helps pinpoint exactly where the routing breaks down in different model types.
Lu: They specifically highlight two distinct architectural forms of this routing failure: an "attention routing failure" in Qwen-family models, where they only attend to image tokens a small percentage of the baseline at the final layer, and a "post-attention integration failure" in InternVL3-8B, which attends well but fails to integrate that attention into its prediction.
Meng: Those two distinct failures give us concrete targets for model development. Knowing whether it's an attention routing issue or an integration issue helps us decide whether we need to modify the attention heads themselves or change how the representations are fed into the final classification layers. That’s practical engineering direction.
Lalam: The paper’s suggestion implies that simply having more visual tokens in the input isn't enough; we need mechanisms that actively guide those tokens toward a specific output layer when they represent a chart, rather than letting them drift or get lost in the general representation space. It’s about intelligent steering.
Conclusion: Tom: So, to wrap up our discussion on "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification," the main conclusion is that chart information is recoverable from intermediate representations but it doesn't reliably reflect in the final verification decision, confirming a routing failure. This confirms what we thought—it’s about flow control, not just raw data intake.
Jane: It really boils down to this: these models are capable of understanding the visual details of a chart when they process it internally, but there's a gap in their internal wiring that prevents that chart-specific information from being utilized during the final prediction step. This is a crucial distinction for understanding model behavior.
Lu: I think the most important implication here is that diagnosing format sensitivity requires looking at internal computation, not just observing task performance metrics on the outside; we need to use these diagnostic analyses when evaluating systems for high-stakes scientific workflows. It’s about understanding the mechanics of the failure.
Meng: I see this as a directive for our teams: before we declare a multimodal system ready for complex applications, we need to run these kinds of internal diagnostics to ensure that evidence from different modalities is actually being routed correctly into the final decision-making units. It’s about building reliability from the inside out.
Lalam: I just want to stress that this research shows us exactly where the weak link is in our current AI systems—it's not always a perceptual failure, but often a structural routing oversight that needs fixing through better architectural design and more targeted training strategies.
Tom: That’s all we have time for today on "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." Thanks for tuning in; we'll be back next time with another fascinating piece of research.
Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin
University of Tokyo · NII LLMC · National Institute of Informatics
cs.CL
Submitted: 2026-06-01
Updated: 2026-10-02
Code: https://github.com/ksunisth/encoded-not-routed
Importance score: 92/100
The gist: Multimodal Large Language Models (LLMs) are increasingly used for scientific peer review, yet they exhibit a significant performance gap when verifying claims supported by charts compared to tables,
Key concepts
- Linear Probing
- This diagnostic technique maps a model's hidden state (its internal representation) to a label prediction using a simple linear classifier. It helps determine if information from different layers of the network is effectively being used to make the final classification decision.
- Attention Analysis
- This method measures how much attention the model pays to image tokens (the visual parts of an input) at different internal layers. It quantifies the total probability mass directed toward visual evidence, showing where the model focuses its visual processing power.
- Routing Failure
- This refers to a problem where relevant information is successfully processed internally but fails to be passed or 'routed' to the specific part of the network responsible for making the final output prediction. In this study, it means chart data is understood but not utilized for verification.
- Mean-Pool Setting
- This probing setting captures task-relevant information that is distributed across many hidden states rather than being concentrated only at the very last token position. It reveals how broadly chart signals are spread throughout the model's processing layers.
Terminology
Summary
Multimodal Large Language Models (LLMs) are increasingly used for scientific peer review, yet they exhibit a significant performance gap when verifying claims supported by charts compared to tables, even when both modalities represent the same underlying data. This paper investigates whether this disparity stems from a failure in extracting visual information or a failure in routing that extracted information to the prediction stage. The core finding is that chart-relevant signal is indeed encoded within the models' intermediate representations but fails to reach the prediction position, reframing the table-chart gap as a routing failure rather than an encoding deficiency.
The Gist
Chart information is encoded in the models’ intermediate representations but does not reach the prediction position, a gap that is absent for tables and holds across all conditions tested.
Experimental Setup and Methodology
The study employed a diagnostic approach to compare how chart and table evidence are processed by three open-weight multimodal LLMs—Qwen2.5-VL-7B, Qwen2.5-VL-32B, and InternVL3-8B—on the SciTabAlign+ benchmark. This dataset pairs 162 unique scientific claims with semantically equivalent table and four chart variants (basic bar chart, symbol bar chart, line chart, and swapped chart), all drawn from the same underlying data. To diagnose the issue of information flow, researchers utilized two primary analytical techniques:
-
Linear Probing: This technique maps a hidden state to a label prediction using a linear classifier. It was applied in two settings:
last-token setting,
where the hidden state is at the final input token position, andmean-pool setting,
which captures task-relevant information distributed across layers rather than concentrated at the final position. -
Attention Analysis: This involved computing the fraction of last-token attention directed to image token positions, normalized by a proportional baseline, to measure total probability mass directed to visual evidence at each layer.
Key Findings from Probing and Attention Analysis
The analysis revealed distinct failure patterns depending on the probing setting:
(a) Last-token probing
For table evidence, the AUROC rises sharply in late layers,
indicating that table information aggregates effectively at the prediction position. In contrast, for chart variants, performance remains near chance across all layers. This suggests a failure of routing where chart information does not concentrate at the final prediction layer.
(b) Mean-pool probing
Under this setting, chart-relevant signal is recoverable from intermediate representations yet is not effectively routed to final verification decisions.
Specifically, meanpool AUROC is higher for chart variants than for table evidence across all models (84–89% vs. 65–70%), the reverse of the last token pattern in Table 1.
This indicates that chart information is more broadly distributed across hidden states but fails to concentrate at the prediction position where table information already aggregates.
Distinguishing Failure Modes Across Model Families
The study identified two architecturally distinct forms of this routing failure:
-
Qwen-family models exhibit an
attention routing failure,
where they largely bypass chart evidence when predicting the claim label. Figure 4 shows that Qwen-family models attend to image tokens only4–11% of the proportional baseline at the final layer.
-
InternVL3-8B shows a
post-attention integration failure.
This model maintains near-proportional aggregate image-token attention (93%) but still fails to outperform its own mean-pool probe on chart variants, suggesting it attends to chart evidence but does not integrate it into the prediction.
Ablation and Further Diagnostic Tests
To further distinguish between perceptual failures and utilization failures, researchers conducted ablation studies:
(Chain-of-Thought Ablation)
Testing explicit chart verbalization before prediction showed that Chain-of-Thought (CoT) prompting worsens macro-F1 for both Qwen models
but improves InternVL3-8B,
suggesting that forcing verbalization of unaddressed chart content yields unreliable descriptions.
(Claim-Only Baseline)
Probes trained only on claim text without image tokens showed a large gap (∆), confirming that the mean-pool probe captures visual information, ruling out claim-text leakage.
This demonstrates that the probe's success is driven by visual information rather than just textual cues.
Conclusion and Ethical Implications
The research concludes that chart information is recoverable from representations but is not reliably reflected in the model’s final verification decision, confirming a routing failure. The two distinct failure modes—attention routing failure in Qwen and post-attention integration failure in InternVL3-8B—underscore that diagnosing format sensitivity requires examining internal computation, not just task performance. This finding serves as a warning that Output evaluation alone cannot surface this distinction,
urging researchers to use diagnostic analyses when evaluating systems for high-stakes scientific workflows.
Improvements for AI systems
Here are specific improvements for AI systems based on the findings in this scientific paper:
-
Improve Scientific Claim Verification Robustness Across Evidence Formats:
-
Develop Format-Agnostic Reasoning Mechanisms:
-
Implement Dynamic Routing and Prediction Gating Strategies:
Detailed Improvements and Capabilities:
-
Improve Scientific Claim Verification Robustness Across Evidence Formats (Addressing the core gap):
-
Develop Format-Agnostic Reasoning Mechanisms (Addressing the
why
of the failure): -
Implement Dynamic Routing and Prediction Gating Strategies (Addressing the
how
to fix it):
Sources
- Qwen2.5-VL Technical Report
- Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering