Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
summary
The gist
Multimodal Large Language Models (LLMs) are increasingly used for scientific peer review, yet they exhibit a significant performance gap when verifying claims supported by charts compared to tables,
In short
Researchers investigated why multimodal AI models struggle to verify scientific claims supported by charts compared to tables. They found that chart information is encoded within the model's internal representations but fails to reach the final decision-making stage. This disparity is not due to poor initial understanding, but rather a failure in routing that directs visual evidence toward the prediction layer.
Key concepts
- Linear Probing
- This diagnostic technique maps a model's hidden state (its internal representation) to a label prediction using a simple linear classifier. It helps determine if information from different layers of the network is effectively being used to make the final classification decision.
- Attention Analysis
- This method measures how much attention the model pays to image tokens (the visual parts of an input) at different internal layers. It quantifies the total probability mass directed toward visual evidence, showing where the model focuses its visual processing power.
- Routing Failure
- This refers to a problem where relevant information is successfully processed internally but fails to be passed or 'routed' to the specific part of the network responsible for making the final output prediction. In this study, it means chart data is understood but not utilized for verification.
- Mean-Pool Setting
- This probing setting captures task-relevant information that is distributed across many hidden states rather than being concentrated only at the very last token position. It reveals how broadly chart signals are spread throughout the model's processing layers.
Terminology used across episodes
This episode discusses
- Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification · Paper Radio
- Qwen2.5-VL Technical Report
- Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
The paper
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification · Read on arXiv
Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin
University of Tokyo · NII LLMC · National Institute of Informatics
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Encoded but Not Routed".
Tom: Multimodal Large Language Models (LLMs) are increasingly used for scientific peer review, yet they exhibit a significant performance gap when verifying claims supported by charts compared to tables,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, what we just discussed was the setup for this study investigating "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." The authors are Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, and Florian Boudin. This paper is really zeroing in on a specific performance disparity where models perform better on tables than charts for scientific claim verification.
Jane: It’s about trying to figure out if the problem is that the AI can't read the chart information at all, or if it can read it but doesn't know how to use that visual data when making its final judgment. The title suggests they are investigating this routing aspect, which is a really interesting angle.
Lu: Exactly! They are testing whether this gap comes from a failure in extracting the visual information itself, or if it stems from a failure in routing that extracted information to the prediction stage, which is what they hypothesize. It’s shifting the focus from perception to utilization.
Meng: A routing failure sounds like a bottleneck in the model's internal architecture, perhaps how different types of evidence—like tables versus charts—are prioritized or passed along through those layers. That makes me think about how we structure our attention mechanisms in these large models.
Lalam: I agree with Lu; if the signal is recoverable from intermediate representations but doesn't reach the prediction position, that's a clear indication that the mechanism responsible for connecting visual evidence to the final decision needs adjustment. It’s about optimizing the flow, not just increasing the raw capacity of each component.
The paper's summary: Tom: Moving on, let's talk about what they actually found in "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." The core finding is quite striking: chart-relevant signal is indeed encoded within the models’ intermediate representations, but it fails to reach the prediction position. This suggests a failure in routing rather than a deficiency in initial encoding of that visual data.
Jane: So, to put it simply, the paper discovered that when you look deep inside the AI's processing layers, you can find evidence of what's happening in a chart, but by the time the model gets to making its final guess about whether a claim is true or false, that specific visual signal is missing from its decision path.
Lu: That confirms their hypothesis regarding routing failure. They used linear probing and attention analysis to show this pattern across three different open-weight multimodal LLMs on the SciTabAlign+ benchmark, which pairs claims with both tables and various chart types. The results were consistent, showing this same pattern regardless of the specific model family they tested.
Meng: That consistency is important for us; it means this isn't just an anomaly in one specific model; it points to a fundamental challenge in how current multimodal AI systems handle visual evidence compared to tabular data structures. It suggests we need a more universal way to integrate different modalities effectively.
Lalam: From my perspective, this summary confirms that the problem isn't that the models are blind to charts; they are blind to routing them correctly into the final decision-making process, which is a much more tractable problem for refinement than trying to teach them how to see better.
The paper's improvements: Tom: Now, let’s discuss what the authors suggest as improvements based on their findings in "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." They are not just pointing out a problem; they are suggesting that we need to look at how models utilize these internal signals more effectively.
Jane: It seems like the paper points toward diagnosing this issue through specific probing techniques, like looking at whether information aggregates sharply at the last token position for tables versus remaining scattered across layers for charts. This helps pinpoint exactly where the routing breaks down in different model types.
Lu: They specifically highlight two distinct architectural forms of this routing failure: an "attention routing failure" in Qwen-family models, where they only attend to image tokens a small percentage of the baseline at the final layer, and a "post-attention integration failure" in InternVL3-8B, which attends well but fails to integrate that attention into its prediction.
Meng: Those two distinct failures give us concrete targets for model development. Knowing whether it's an attention routing issue or an integration issue helps us decide whether we need to modify the attention heads themselves or change how the representations are fed into the final classification layers. That’s practical engineering direction.
Lalam: The paper’s suggestion implies that simply having more visual tokens in the input isn't enough; we need mechanisms that actively guide those tokens toward a specific output layer when they represent a chart, rather than letting them drift or get lost in the general representation space. It’s about intelligent steering.
Conclusion: Tom: So, to wrap up our discussion on "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification," the main conclusion is that chart information is recoverable from intermediate representations but it doesn't reliably reflect in the final verification decision, confirming a routing failure. This confirms what we thought—it’s about flow control, not just raw data intake.
Jane: It really boils down to this: these models are capable of understanding the visual details of a chart when they process it internally, but there's a gap in their internal wiring that prevents that chart-specific information from being utilized during the final prediction step. This is a crucial distinction for understanding model behavior.
Lu: I think the most important implication here is that diagnosing format sensitivity requires looking at internal computation, not just observing task performance metrics on the outside; we need to use these diagnostic analyses when evaluating systems for high-stakes scientific workflows. It’s about understanding the mechanics of the failure.
Meng: I see this as a directive for our teams: before we declare a multimodal system ready for complex applications, we need to run these kinds of internal diagnostics to ensure that evidence from different modalities is actually being routed correctly into the final decision-making units. It’s about building reliability from the inside out.
Lalam: I just want to stress that this research shows us exactly where the weak link is in our current AI systems—it's not always a perceptual failure, but often a structural routing oversight that needs fixing through better architectural design and more targeted training strategies.
Tom: That’s all we have time for today on "Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification." Thanks for tuning in; we'll be back next time with another fascinating piece of research.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization