Using OCR Heads to Verbalize Image Semantics
cs.CV, cs.AI, cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 21 pages, 22 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR).
Terminology
Abstract
How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word "bike" causes Qwen3-VL-8B to output "bike," but pointing them at a bird wing causes the model to output the token "feathers." We collapse these heads' attention weights into a single verbalization lens transformation that reveals interpretable semantic features in hidden states across all layers. When combined with projection to vocabulary space, we can obtain interpretable labels starting from layer 0, showing that image representations are in fact aligned with language in early layers. We find that we can also use the inverse of this transformation to edit non-word concepts, e.g., replacing a tractor with a revolver in a naturalistic image, providing causal evidence that this subspace is useful for more than just OCR. Our results are an example of how the study of specific mechanisms can shed light on broader interpretability problems.
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Gaze Heads: How VLMs Look at What They Describe
- How Do Vision-Language Models Process Conflicting Information Across Modalities?
- The Platonic Representation Hypothesis
- Decomposing Query-Key Feature Interactions Using Contrastive Covariances
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- ClipCap: CLIP Prefix for Image Captioning
- Towards Interpreting Visual Information Processing in Vision-Language Models
- Interpreting the linear structure of vision-language model embedding spaces
- ImageNet Large Scale Visual Recognition Challenge
- VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
- Function Vectors in Large Language Models
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models