Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
cs.CV, cs.AI, cs.CL, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-29
Comments: 40 pages, 9 figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output.
Terminology
Abstract
When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages 0.740 versus 0.432 native macro-F1, while residual reconstruction reaches 0.486, whereas Gemma improves from 0.532 to 0.714. These differences reflect supervised accessibility rather than a pre-existing, native decision rule, and the most influential token role depends on the task. Under the evaluated score scales, Qwen silent-feature ablation is 24-63 times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is 16-140 times more output-sensitive. Calibration-only routing recovers 93.3 % of the mean gap, and probe-distilled LoRA improves native predictions, although shared multi-task adaptation causes negative transfer. A case study of Gemma-3-12B on Facebook Hateful Memes finds a distributed rank-32 image-prompt interaction, reaching 0.756 versus 0.685 native macro-F1. Robustness controls show that the signal extends beyond English, is not explained solely by accompanying OCR, and depends on paired visual evidence. Thus, routing, rather than representation alone, is a recurring bottleneck in harmful meme classification.
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Discovering Latent Knowledge in Language Models Without Supervision
- Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- Distilling the Knowledge in a Neural Network
- LoRA: Low-Rank Adaptation of Large Language Models
- Hadamard Product for Low-rank Bilinear Pooling
- Visual Credit Audit for Multimodal Spatial Reasoning
- SAE-V: Interpreting Multimodal Models for Enhanced Alignment
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Detecting Harmful Memes and Their Targets
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models