Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

arXiv:2605.13737 · cs.AI, cs.CL · Submitted 2026-05-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Senses Wide Shut".

Jane: When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re jumping into "Senses Wide Shut: A Representation–Action Gap in Omnimodal LLMs," and it really zeroes in on that tricky area where an AI might accept a statement that clashes with its actual sensory input. We’re looking at how the team separates whether the problem lies with seeing or hearing versus what the AI does next.

Jane: That’s right, Tom; it focuses on that fundamental question: does a model fail because it can't see something wrong, or because it fails to translate that perceived error into an output rejection? They use a new setup called IMAVB to test this specific dissociation.

Lu: From my perspective, Lu thinks the design of IMAVB is really clever because it forces us to separate the detection of a conflict from just regular multimodal comprehension, which existing benchmarks haven't done well. It’s designed to test if the model just blindly trusts the text or if it actually checks its senses.

Meng: I’m curious about how this separation works practically; does it mean we can isolate a vision failure from an audio failure when we look at these results? We need to know how granular this conflict detection is for real-world deployment.

Lalam: I think Lalam would focus on the underlying mechanism, and the paper shows that even when models don't reject the false claim in their final answer, their hidden states reliably hold this mismatch information. That means there’s a signal in the model’s internal representation that we are currently missing.

Tom: If it’s encoded in those hidden states but not showing up as a rejection, that points to a translation issue between what the model sees internally and what it actually outputs. Jane, can you explain the representation-action gap in simpler terms for our listeners?

Jane: Absolutely; the paper explains that this gap means the model’s internal understanding of a mismatch isn't translating into a real rejection when it speaks. It’s like having a warning light on inside the machine, but the dashboard doesn't actually show an alert based on that light.

Lu: I think this is fascinating because it suggests the bottleneck isn't necessarily perception itself, but rather how that perception gets translated into a decision or action. It shifts the focus from just improving visual processing to improving the internal mapping process.

Title and authors: Meng: From an engineering standpoint, if this gap exists across different models, it means we need to build in some kind of feedback loop during inference to force that internal mismatch signal into a rejection behavior. How do you suggest we engineer that feedback loop?

Lalam: Lalam would point out the diagnostic tools they used, like linear probing, which showed the standard versus misleading distinction is linearly decodable from hidden states. That’s a concrete way to see where that signal originates within the model's structure.

Tom: That’s heavy, Lalam; so they aren't just guessing where the problem is; they are actually mapping it inside the AI architecture. Jane, what about those specific improvements they suggest for fixing this gap? What are we actually supposed to do next?

Jane: The paper proposes things like Probe-Guided Logit Adjustment, or PGLA, which involves injecting that encoded mismatch signal back into the decoding process. It basically trains a small part of the model to predict the misleading split probability based on what it already knows.

Lu: I think PGLA is interesting because it treats the mismatch as an actionable signal, which is a big step toward making the system more robust against these kinds of textual contradictions. It moves the problem from passive encoding to active rejection.

Meng: If PGLA gives a gain of plus fifteen percentage points across eight models, that’s a very tangible improvement we can measure in a lab setting. My concern is making sure this adjustment doesn't introduce new biases when we apply it to different types of audio or vision tasks.

Lalam: I think the findings on modality asymmetry are key here; the paper noted that vision-grounded probes performed better than audio-grounded ones, and that the gap for audio misleading questions was actually larger. That tells us we might need to apply different fixes depending on which sense is involved.

Tom: So, we have a clear direction: not just better perception, but better translation from perception to action through these kinds of adjustments. Jane, what do you think the long-term impact of this research is for general AI development?

Jane: I see it as moving us toward systems that are fundamentally more trustworthy because they can detect when their internal understanding doesn't match the external reality presented in a question. It helps build confidence in deploying multimodal AI where sensory grounding is crucial.

Title and authors: Lu: The implications for future work, as suggested by the authors, point toward training objectives that explicitly tie what the model says to what it has already detected internally. That suggests a fundamental shift in how we design the learning process itself.

Meng: I'm thinking about the practical side of that; if we can create objectives that enforce this alignment, it could lead to more reliable agent control systems, perhaps for things like robotic health attendants where context and reality matter.

Lalam: Lalam would emphasize the need for continuous self-improvement through diagnostic frameworks, checking those hidden states periodically to ensure that signal doesn't decay before it reaches the output layer. That’s how we keep chasing that perfect grounding.

Tom: It sounds like a really solid direction for the next generation of models, focusing on making sure the internal processing actually drives the final behavior. Jane, before we wrap up this discussion on "Senses Wide Shut," what's your final thought on what this paper really means for us right now?

Jane: I think it means we need to stop treating perception and action as separate boxes and start designing systems where they are inherently tied together from the start. It’s about making sure that when an AI interacts with the world, its internal reality matches what it communicates externally.

Lu: The whole IMAVB setup is a really important contribution because it provides a rigorous, multi-faceted way to test this specific kind of failure mode across different modalities and conditions. It gives us the tools we need to diagnose these kinds of deep structural issues.

Meng: From my side, I'm looking at how this diagnostic framework could be used to audit existing systems for potential grounding weaknesses before they go into production. That kind of proactive safety check is what engineers need.

Lalam: Lalam would conclude that the core message of "Senses Wide Shut" is that detection happens internally, but the system needs a better way to communicate that detection outwardly. That translation gap is the central challenge we need to solve.

Tom: So, we've seen how this paper uses IMAVB and PGLA to diagnose a specific failure mode, showing that the problem is often in the translation step rather than just raw perception. Jane, Lu, Meng, Lalam—that was some deep stuff on the representation-action gap.

Title and authors: Jane: That’s a great summary of how we diagnose this issue using the IMAVB setup and PGLA techniques. It clearly shows that the internal signal is real, even if it doesn't immediately manifest as a rejection.

Lu: And I think the diagnostic tools they used—linear probing and logit lens analysis—are really insightful because they let them actually map where that translation failure happens within the AI's structure. They didn't just guess; they found the exact spot in the hidden states where the signal gets lost.

Meng: If we’re looking at practical fixes, I see PGLA as a repeatable method to improve rejection rates without having to completely redesign the model architecture. That sounds like something we could try implementing in our own model fine-tuning pipelines.

Lalam: I agree with Meng; if we can successfully inject that encoded signal, it means we’re creating a mechanism where the AI actively checks its internal conflict before giving an answer, which is a significant step toward more reliable behavior.

Tom: So, to wrap up our discussion on "Senses Wide Shut," we’ve seen how this study exposes a critical failure mode where an AI accepts a false premise because it misses the contradiction between what it sees and what it hears. Jane, Lu, Meng, Lalam—that was some deep stuff on the representation-action gap.

Jane: It really boils down to this idea that the model’s internal understanding of a sensory mismatch doesn't actually translate into a correct rejection in its final output. We need to focus on fixing that translation step.

Lu: I think the most significant implication is that we are moving beyond just improving raw data ingestion; we have to focus on designing better mechanisms for how those sensory inputs get mapped onto decision-making pathways within the AI structure itself.

Meng: From a practical standpoint, this means any future deployment of multimodal AI needs to incorporate these internal checks, otherwise, it’s just a sophisticated guessing machine that might be dangerously wrong. We need to make sure we’re engineering for reliability at every layer of decision-making.

Lalam: For me, the real impact is on building public trust; if we can create systems that are internally consistent and refuse to lie based on sensory contradictions, it sets a new standard for how dependable AI can be in our world.

The paper's summary: Tom: Now that we’ve established the context of the IMAVB benchmark, let’s get into what "Senses Wide Shut: A Representation–Action Gap in Omnimodal LLMs" actually tells us about the core issue. Essentially, it shows that even when an omnimodal AI has all its senses working—vision and audio—it can still completely miss a contradiction between what it sees and what it hears because of a translation problem.

Jane: Exactly, Tom; the paper demonstrates that there's this hidden signal inside the model's brain indicating a mismatch, but that signal never actually gets translated into the correct response when the AI speaks or acts on it. It’s like having a warning light on in your system, but you never actually look at it on your dashboard.

Lu: And what I find really wild is that this isn't just about perception failing; it points to a breakdown in the process of taking sensory input and turning it into an output decision. The authors call this the representation-action gap, suggesting the bottleneck lies in that bridge between what's processed internally and how it behaves externally.

Meng: From an engineering standpoint, I see that as incredibly frustrating because you can have a super sophisticated perception engine, but if the action layer can't interpret that perception correctly, you’re stuck with flawed reasoning. The paper suggests we might need to build in ways to actively force that internal mismatch signal into the final decision-making process during inference.

Lalam: I think this has huge implications for how we trust multimodal systems; if the AI doesn't show it recognizes a conflict, it might be acting on false information without realizing its foundation is shaky. This could dramatically affect how we deploy these models in safety-critical applications where sensory grounding is vital.

Tom: Right, Lalam, that’s huge because it moves the focus from just making the AI *see* better to making sure it *connects* what it sees to what it says. It’s not just a perception problem anymore; it’s a communication problem inside the model itself.

Jane: And those diagnostic tools they used, like linear probing and logit lens analysis, were really insightful because they let them actually map where that translation failure happens within the AI's structure. They didn't just guess; they found the exact spot in the hidden states where the signal gets lost.

Lu: The idea of a "translation-bottleneck regime" versus an "unembedding-misaligned regime" suggests different ways this gap can form, which opens up so many creative avenues for future research on how to fix these structural weaknesses. We could build entire new architectures designed specifically to solve that translation issue.

Meng: If we're looking at practical fixes, the paper proposes things like PGLA, which tries to re-inject that mismatch signal back into the decoding process to force a rejection behavior. That sounds like something we could try implementing in our own model fine-tuning pipelines.

Lalam: I agree with Meng; if we can successfully inject that encoded signal, it means we’re creating a mechanism where the AI actively checks its internal conflict before giving an answer, which is a significant step toward more reliable behavior.

Tom: So, the core message for us is this: the future of omnimodal AI isn't just about increasing the raw processing power of vision and audio encoders; it’s about figuring out how to build smarter translation layers that ensure what we see aligns with what we say.

Jane: That’s a wonderful way to put it, Tom; it shifts our focus from simply adding more sensors to ensuring those sensors are perfectly synced in the model's mind. It makes the concept of "grounding" much deeper than just matching images to text.

Lu: And I think this opens up incredible creative possibilities for how AI can reason about complex, multi-sensory scenes where contradictions are common and subtle. Imagine an AI that can tell you *why* it doesn't agree with a premise based on a hidden visual detail it caught but failed to fully process in its initial output.

Meng: I think the immediate implication is that we need better auditing tools for deployed models, because if we can’t diagnose this gap internally, how do we know when an AI is being misled by its own flawed internal logic?

Lalam: I think the cultural impact is tied to trust; if these systems are reliably grounded, people will feel much more secure interacting with complex digital environments built around these AIs. It builds a foundation for more dependable technology in everyday life.

The paper's improvements: Tom: So, we’ve seen how the authors diagnosed this internal translation issue using tools like linear probing, and now they’re proposing specific fixes to bridge that gap in their paper "Senses Wide Shut." They aren't just pointing out a problem; they are giving us actionable methods to fix it.

Jane: That's right, Tom; they suggest techniques like PGLA, which essentially re-inject the detected mismatch signal directly into the AI’s decoding process during inference. It’s a concrete method for intervention.

Lu: That’s really exciting because it moves the solution from passive detection to active correction; it turns the internal knowledge of a conflict into a concrete rejection behavior in real-time. It suggests that we can train models not just to recognize what's wrong, but to actively fight against that specific type of error.

Meng: As an engineer, I’m interested in PGLA because if it consistently yields a gain across different models, it shows we have a repeatable way to improve rejection rates without having to completely redesign the entire model architecture. It’s a surgical intervention.

Lalam: I think that ability to actively correct the output is what makes this paper so impactful for cultural standards of AI; if we can build systems that self-correct based on internal conflict, it establishes a much higher bar for what we consider trustworthy multimodal interaction. It moves us toward more dependable tools in society.

Tom: Exactly, Lalam, and I think the paper’s suggestion to implement modality-specific thresholds is key here because it acknowledges that vision and audio aren't equally reliable for every kind of claim. We can tailor the fix based on whether the AI is failing on visual or auditory input.

Jane: And those adaptive rejection strategies are brilliant because they show that we don't need a one-size-fits-all solution; we need systems that understand the inherent weaknesses of different sensory modalities when handling conflicting information. It’s about smart adaptation rather than brute force correction.

Lu: The authors also touched on the idea of training objectives like honesty-targeted alignment, which is a big theoretical step toward making the model's internal representation perfectly match its final utterance. That level of alignment is what we’ve been chasing for a long time in theoretical AI research.

Meng: If we focus our engineering efforts on those types of contrastive grounding losses they mentioned, it gives us a clear training target to push the model away from that representation-action gap during the learning phase itself, which is much more efficient than post-hoc adjustments like PGLA.

Lalam: I think focusing on those objective alignments means we are building fundamentally more robust AI that learns to be truthful about its own understanding of the world, which is essential for building a culture where we can rely on these technologies.

Tom: So, to sum up the improvements: we’re moving toward systems that use internal conflict signals to actively adjust their output in a modality-aware way, rather than just hoping the initial perception was correct. It’s an exciting direction for how we design next-generation models.

Jane: That’s a great way to put it, Tom; it emphasizes that the future isn't just about bigger models but about smarter internal feedback loops that ensure what the AI communicates matches what it genuinely perceives.

Lu: And this whole line of research, looking at how to close those gaps in translation between perception and action, could fundamentally reshape how we approach building truly integrated multimodal intelligence.

Meng: I’m just waiting to see how scalable these correction mechanisms are when we move from 7B models to much larger architectures; that's the main practical hurdle for me right now.

Lalam: But even with scaling challenges, the potential for creating AI that is inherently more honest and trustworthy across different sensory inputs is something we need to be excited about.

Conclusion: Tom: So, to wrap up our discussion on "Senses Wide Shut: A Representation–Action Gap in Omnimodal LLMs," we’ve seen how this study exposes a critical failure mode where an AI accepts a false premise because it misses the contradiction between what it sees and what it hears.

Jane: It really boils down to this idea that the model’s internal understanding of a sensory mismatch doesn't actually translate into a correct rejection in its final output, which is something we need to fix.

Lu: I think the most significant implication is that we are moving beyond just improving raw data ingestion; we have to focus on designing better mechanisms for how those sensory inputs get mapped onto decision-making pathways within the AI structure itself.

Meng: From a practical standpoint, this means any future deployment of multimodal AI needs to incorporate these internal checks, otherwise, it’s just a sophisticated guessing machine that might be dangerously wrong. We need to make sure we’re engineering for reliability at every layer of decision-making.

Lalam: For me, the real impact is on building public trust; if we can create systems that are internally consistent and refuse to lie based on sensory contradictions, it sets a new standard for how dependable AI can be in our world.

Tom: Exactly, Lalam; this research gives us the map to fix that translation gap using things like PGLA and modality-specific adjustments. It’s a huge step toward making these systems more accountable.

Jane: We're leaving this segment with the understanding that the next frontier isn't just about how much data an AI can process, but how intelligently it processes and connects what it experiences across different senses.

Lu: That opens up so many creative avenues for building entirely new architectures that prioritize this cross-modal consistency from the very beginning of the design process.

Meng: I’m looking forward to seeing more work on how these diagnostic tools can be integrated into automated testing pipelines for safety checks in production environments.

Lalam: And I think this paper, "Senses Wide Shut," is a crucial piece of the puzzle for developing a culture where AI is not just powerful, but fundamentally reliable.

Nanyang Technological University LMMs-Lab Team Johns Hopkins University

cs.AI, cs.CL

Submitted: 2026-05-13

Updated: 2026-10-01

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 90/100

The gist: When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? The core finding is that

Key concepts

IMAVB Benchmark
This is a curated dataset of 500 movie clips designed to test multimodal models. It features a 2x2 design crossing vision and audio modalities with standard and misleading premise conditions, allowing researchers to measure conflict detection separately from general understanding.
Representation-Action Gap
This gap describes the disconnect where a model's internal hidden states successfully encode conflicts between text premises and sensory evidence. However, this encoded signal does not reliably translate into the desired external action, such as rejecting a false claim in the model's final response.
Translation-Bottleneck Regime
This is an analysis of how information flows through the model during processing. In this regime, the correct answer is visible within the internal representation matrix but decays before it reaches the final output layer, indicating that translation between detection and decision-making is slow or incomplete.

Terminology

Summary

When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? The core finding is that hidden states reliably encode these premise–perception mismatches even when models almost never reject the false claim in their outputs, suggesting the bottleneck for omnimodal grounding lies in translation rather than perception.

The Benchmark and Design

The study introduces IMAVB, a curated 500-clip benchmark of long-form movies with a 2×2 design crossing target modality (vision, audio) and premise condition (standard, misleading). This design allows for the measurement of conflict detection separately from ordinary multimodal comprehension. The benchmark is constructed to meet five properties that existing benchmarks lack simultaneously: all modalities must remain intact, false premises must be implicit rather than explicit verification prompts, modality targeting must be surgical to measure vision-grounded and audio-grounded failures separately, stimuli must extend over time for cross-modal tracking, and symmetric standard controls must exist. The dataset consists of 500 clips (1–5 min) from three sources: @BingeSociety, @BoxofficeMoviesScenes, and Condensed Movies.

The Representation–Action Gap

The research documents a Representation–Action Gap, where hidden states reliably encode premise–perception mismatches even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions. This gap is characterized as modality-asymmetric (audio grounding underperforms vision) and prompt-resistant across seven variants.

Diagnostic Interpretability Tools

To investigate the source of the gap, the authors employ two complementary interpretability tools: linear probing and logit lens projection. Linear probing shows that the standard /misleading distinction is linearly decodable from hidden states, with modality-specific probes revealing a 9–15pp internal gap across the six standard 7B models in residualized accuracy, indicating the signal arises from multimodal processing. Logit lens analysis reveals two regimes: a translation-bottleneck regime where the correct answer is readable through the unembedding matrix mid-stack but decays before output, and an unembedding-misaligned regime where the misleading signal is decodable by a learned probe but never aligns with the unembedding’s column space.

The Probe-Guided Logit Adjustment (PGLA)

As an initial diagnostic intervention, the authors propose PGLA, which re-inject[s] the encoded mismatch signal into decoding and consistently improves rejection behavior. This involves extracting hidden states at a peak probe layer and training a two-layer MLP to predict the misleading split probability. The confidence-gated adjustment modifies logits for rejection options E and F using an equation that incorporates the probe confidence, a gap-adaptive term, and a debiasing term to correct for positional bias. PGLA yields a mean gain of +15.0pp across all eight open-source models, providing diagnostic evidence that the encoded signal is actionable.

Key Findings on Asymmetry and Failure Modes

The analysis reveals significant modality asymmetry: vision-grounded probes perform better than audio-grounded ones, and the behavioral gap for audio misleading questions is larger. Furthermore, Qwen3-Omni exhibits an outlier behavior where both standard and misleading accuracy improve under PGLA, suggesting its mechanism differs from others. The research concludes that the failure lies in translation: hidden states contain signal sufficient to improve premise rejection when fed back to the output, yet the default output distribution does not reflect this encoding. This suggests future work should focus on training objectives that tie what the model says to what it has already detected internally.

Conclusion and Impact

The study concludes that hidden states pick up on mismatch reliably, but this signal does not reach the output. The IMAVB benchmark is proposed for integration with lmms-eval, and all code will be released publicly. The findings diagnose a safety-relevant failure mode in omnimodal LLMs: the systematic inability to detect when textual premises contradict sensory evidence, which enables developers to build more trustworthy multimodal systems.

The gist: Hidden states reliably encode premise–perception mismatches even when models almost never reject the false claim in their outputs, suggesting the bottleneck for omnimodal grounding lies in translation rather than perception.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the findings in this research, along with what those improved systems will be capable of:

  1. Do not rely solely on textual premises for grounding; implement a multi-modal verification system that cross-references textual claims against simultaneous sensory input (video and audio). This system will detect silent compliance where a model accepts a false premise because it fails to notice the sensory contradiction.

  2. Develop an internal mechanism to explicitly encode and propagate detected perception–action mismatches. Instead of simply outputting text, the model should generate a signal (as suggested by the Representation–Action Gap) that is fed back into its decoding process via an inference-time adjustment (like PGLA). This will allow the system to actively reject misleading questions based on internal conflict, rather than relying on external prompt engineering.

  3. Implement modality-specific grounding thresholds and adaptive rejection strategies. The system should learn that audio grounding may be inherently less reliable than vision grounding for certain types of claims (modality asymmetry), allowing it to adjust its confidence or rejection behavior dynamically when processing conflicting signals from different sensory channels.

  4. Incorporate truthful explanation generation as a core objective, particularly for binary verification tasks. The system should be trained not only to predict the correct answer (TRUE/FALSE) but also to generate a justification that accurately reflects the ground truth of its own prediction, ensuring that rejections are grounded in perceptual reality rather than heuristic shortcuts.

  5. Improve robustness against prompt variations and positional bias through systematic evaluation and mitigation techniques (like option shuffling). The system's internal reasoning should be invariant to the order of options provided, ensuring that detection capabilities are not artifacts of how the question is phrased.

  6. Utilize a diagnostic framework for continuous self-improvement. By periodically running hidden-state probes (linear probing) on its internal representations and checking for signal decay toward the output layer, developers can identify where the representation–action gap occurs and target training objectives—such as contrastive grounding losses or honesty-targeted alignment—to fix the translation bottleneck, rather than simply scaling up model size.

These improvements will result in AI systems that are significantly more trustworthy, resilient to adversarial inputs designed to exploit textual biases, and capable of performing reliable reasoning on complex audiovisual tasks where sensory evidence must be integrated coherently.

Abstract

When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation-Action Gap: hidden states reliably encode premise-perception mismatches even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (audio grounding underperforms vision) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior. Together, these results suggest the bottleneck for omnimodal grounding lies in translation, not perception.

Sources

Related papers