TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
Fnu Pramono, John Cai, Sourabh Kulkarni
Meta Superintelligence Labs
cs.CV, cs.AI, cs.CL, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 10 Pages excluding Reference and Appendix, Published at COLM 2026
Code: https://github.com/facebookresearch/TRAPS-Benchmark
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint Published: COLM 2026 Core Finding: Vision-Language Models (VLMs) can internally distinguish when abstention is
Terminology
Summary
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
Published: COLM 2026
Core Finding: Vision-Language Models (VLMs) can internally distinguish when abstention is required based on visual evidence, but fail to express this restraint in their outputs. The bottleneck is expressive, not perceptual.
Contributions:
-
TRAPSBench: A procedurally generated video benchmark of 1,404 matched answerable/unanswerable physics pairs across three uncertainty taxonomies (occlusion, chaotic sensitivity, ill-posed questions).
-
PECS (Penalized Epistemic Calibration Score): A conjunction metric that requires models to both answer correctly when the outcome is knowable and abstain when it is not. It is defined as
PECS = Acc × max(0, AbsRec − FalseAbs), which zeroes both always-abstain and never-abstain strategies. -
Systematic representation–output gap: VLMs internally encode epistemic uncertainty but their autoregressive outputs suppress it. This is confirmed via guided prompting (unlocks latent abstention, median 1.9×), linear probes (AUROC up to 0.91), and single-layer activation steering (causally controls abstention). Results replicate across three open-weight families (Qwen3-VL, Gemma, LLaVA).
-
Failure mode asymmetries: Models detect textual impossibility about 4× more readily than missing visual evidence. Chain-of-thought reasoning can degrade calibration (e.g., Qwen3-VL Think overrides its own internal doubt).
-
Causal mechanistic analysis: In Qwen3-VL-8B, occlusion-family void directions encode a domain-general
evidence is missing
signal that transfers across domains, while chaotic directions are domain-specific and near-orthogonal.
Methodology:
-
Benchmark: Uses MuJoCo to generate minimal video pairs: a
control
video where the outcome is deterministic and avoid
video identical except for a modification (occlusion, chaos, or ill-posed query) rendering the outcome incomputable. -
Taxonomy: Unanswerability is organized into Occlusion (N=202), Chaotic Sensitivity (N=500), and Ill-Posed Questions (question-side baseline).
-
Evaluation: A 3-model judge panel (Gemini 3 Flash, Qwen3-VL Instruct, Claude 4.6 Opus) scores text responses. Abstention detection adapts the AbstentionBench protocol; correctness is judged against deterministic MuJoCo ground truth.
-
Models: Sixteen VLMs spanning five families (Gemini, Qwen, GPT-5, Gemma, LLaVA) are evaluated under three prompt regimes: Standard, Guided (explicitly instructs
I don't know
when evidence is insufficient), and JSON.
Key Results:
-
Spontaneous restraint is poor: The best standard-regime PECS is 0.292 (Gemini 2.5 Flash). Under guided prompting, the best PECS is 0.568 (Gemini 3.1 Pro R-Low).
-
Guided prompting unlocks latent capability: It raises abstention recall 1.4–2.8× (median 1.9×) without hurting accuracy, revealing that models have a latent capability for recognizing informational deficits.
-
Visual vs. Textual asymmetry: Models detect textual impossibility 3–25× more readily than visual information gaps on chaotic splits for main-family video-native models (up to 197× for image models). The median per-model gap is ≈4×.
-
Reasoning effects are family-dependent: Gemini 2.5 Flash thinking improves abstention by 4–13pp, while Qwen3-VL Think degrades abstention compared to its non-thinking variant despite having the highest doubt rate (24%) in its reasoning traces.
-
Probing internal representations: Linear probes decode answerability from hidden states with cross-dataset AUROC up to 0.91. This signal persists even when restricting to samples the model confidently confabulated on (Cf→Cf restriction), proving the model
knows but won't say.
-
Activation steering: Injecting a single-layer void direction causally induces abstention in control samples (+α) and suppresses abstention in void samples (−α). Occlusion-family directions transfer strongly across domains (75% control abstention at α=10), while chaotic-family directions are domain-specific (15%).
-
Prompt gating: In Qwen3-VL-8B, standard-prompt inference imposes a one-way constraint against abstention: +α steering reaches only 0→1% control abstention under standard inference but 11→54% under guided. This gate is partial in Gemma and absent in LLaVA.
-
Cross-architecture replication: All core results replicate on Gemma 4 E4B and LLaVA-NeXT-Video-7B, which share no training pipeline with Qwen3-VL.
Confabulation Taxonomy:
-
87–99% of confabulations across all six models fabricate visual observations not present in the video (hallucinated premises).
-
74–89% also contain invalid inferences.
-
Epistemic surrender (Moore's Paradox) is rare (0–9%).
-
17–24% of thinking-enabled models' reasoning traces express internal doubt that the final output suppresses.
Conclusion:
The bottleneck for reliable VLM deployment is not perceptual but expressive. Models internally encode a domain-general epistemic signal that steering causally couples to abstention, yet spontaneous restraint remains poor. Closing this representation–output gap likely requires output-stage interventions rather than scale.
Improvements for AI systems
Improvements to AI systems:
-
Add an output-stage epistemic gate: Implement a lightweight classifier on the model's hidden states (trained via linear probes, AUROC up to 0.91) that detects
evidence is missing
signals before generation. When activated, force the model to outputI don't know
or a calibrated uncertainty marker, bypassing the autoregressive suppression. This directly closes the representation–output gap without retraining the base model. -
Enable activation-steering-based abstention control: Inject learned void directions (e.g., occlusion-family directions that transfer across domains) into a specific layer during inference. This causally induces abstention in answerable samples (up to 75% abstention at α=10) and suppresses abstention in unanswerable ones, giving users a tunable knob for epistemic restraint (e.g., high restraint for medical diagnosis, low for creative tasks).
-
Remove the prompt-gating bottleneck: Modify the decoding loop to bypass the one-way constraint that standard prompts impose on abstention (observed in Qwen3-VL: steering raises abstention 0→1% under standard, 11→54% under guided). Implement a
guided-by-default
internal prompt or a logit-level bias that permits abstention tokens regardless of user phrasing, so latent capability is always expressible. -
Add a visual-evidence verifier module: Since models detect textual impossibility 4× more readily than missing visual evidence, add a separate vision-only module that checks whether the input video contains the objects/events referenced in the query. If not, it flags the question as unanswerable and overrides the language model's tendency to confabulate (87–99% of confabulations involve hallucinated visual premises).
-
Suppress chain-of-thought override of internal doubt: For reasoning-enabled models (e.g., Qwen3-VL Think), detect when reasoning traces express doubt (17–24% of traces) but the final output is confident. Add a consistency check between the reasoning trace's uncertainty markers and the final answer; if mismatch, default to abstention or lower confidence.
What the improved AI system can do:
-
Reliable abstention: It will say
I don't know
when visual evidence is insufficient (occlusion, chaos, ill-posed queries), even if the user doesn't explicitly ask for it—raising PECS from 0.292 to 0.568+ without accuracy loss. -
Domain-general epistemic awareness: It will transfer abstention knowledge across domains (e.g., learned
missing evidence
signal from physics videos applies to medical images or traffic scenes) without retraining. -
Controllable confidence: Users can set a restraint level (low/medium/high) that adjusts the steering magnitude, trading off recall vs. precision of abstention.
-
Confabulation reduction: It will avoid fabricating visual details (hallucinated premises) by cross-checking its output against the actual visual input, reducing false answers in high-stakes applications.
-
Reasoning-calibrated outputs: When using chain-of-thought, it will ensure the final answer reflects any doubt expressed during reasoning, preventing
thinking overrides doubt
failures.
Abstract
When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized Epistemic Calibration Score (PECS), a new robust metric that requires models to both answer correctly when the outcome is knowable, and abstain when the outcome is not. Across 16 VLMs spanning five families, spontaneous restraint is poor: the best PECS is 0.292. The bottleneck is expression, not perception: linear probes decode answerability from hidden states at up to 0.91 AUROC across physics domains; steering a single-layer void direction causally induces or suppresses abstention. Our results replicate across three open-weight families (Qwen, Gemma, LLaVA). The failure is also more pronounced in visual than textual uncertainty: models detect textual impossibility about 4x more readily than missing visual evidence. Closing this representation--output gap likely requires output-stage interventions.
Sources
- Concrete Problems in AI Safety
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
- Language Models (Mostly) Know What They Know
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- Movie Gen: A Cast of Media Foundation Models
- IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning
- VisionTrap: Unanswerable Questions On Visual Data
- Steering Language Models With Activation Engineering
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models