GUIDE: Guiding Internal Evidence with Language Instructions
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Journal ref: EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large multimodal models follow instructions about what to generate, but not necessarily about what evidence to rely on.
Terminology
Abstract
Large multimodal models follow instructions about what to generate, but not necessarily about what evidence to rely on. Hence, models may continue to depend on shortcut-associated cues even when instructions suggest otherwise. We introduce GUIDE, a framework for controlling internal evidence usage through language instructions. GUIDE combines grouped parameter-efficient adaptation with instruction-conditioned gating to modulate multimodal evidence pathways during reasoning and generation. We further introduce a pathway-level evaluation framework that characterizes instruction-conditioned evidence modulation through reliance sensitivity, controlled perturbation analysis, pathway modulation, and autoregressive decoding dynamics. Across multimodal reasoning, classification, and generation, GUIDE induces structured and instruction-aligned redistribution of evidence reliance while largely preserving task behavior. Experiments on GQA, TextVQA, MM-IMDb, CREMA-D, RAVDESS, and Flickr30K show that GUIDE improves robustness under targeted evidence perturbations and enables controllable modulation across diverse multimodal settings. This suggests that multimodal instruction following can extend beyond output control toward regulating how different evidence sources contribute to model predictions.
Sources
- Gated Multimodal Units for Information Fusion
- Transformer Interpretability Beyond Attention Visualization
- Visual Instruction Inversion: Image Editing via Visual Prompting
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- Steering Language Models With Activation Engineering
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering