Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
cs.CL, cs.AI, cs.CV, cs.DL
Submitted: 2026-05-26
Updated: 2026-09-07
Code: https://github.com/tesseract-ocr/tessdata
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on language priors.
Terminology
Abstract
Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on language priors. Comparing open-weight VLMs with traditional OCR baselines on low-resource Ancient Greek critical editions, we show that VLM errors often remain fluent even when wrong, producing plausible Greek substitutions where traditional engines produce local recognition noise. To analyze visual evidence during decoding, we introduce controlled image perturbations and token-level grounding measures based on conditional versus image-free decoding distributions. Under character-level perturbations, VLMs diverge sharply from the perturbed ground truth while traditional OCR remains comparatively faithful; however, token-level analysis shows that prior reliance is model-specific: in an OCR-specialist model, fluent lexical errors are produced with little reliance on the image, whereas general-purpose VLMs remain conditioned on the visual input even when wrong. Decode-time interventions fail to reliably restore grounding, while post-OCR language-model correction improves several systems only by repairing text after generation. Our results extend prior evidence of OCR language-prior reliance to low-resource historical documents and a broader set of models, showing that fluent output is not necessarily visually grounded and motivating interpretability-driven evaluation beyond aggregate accuracy.
Sources
- Pixtral 12B
- Structure-Aware Text Recognition for Ancient Greek Critical Editions
- Qwen3-VL Technical Report
- Geometric Risk Control for Vision-Language Model OCR
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
- GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
- Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
- Logios : An open source Greek Polytonic Optical Character Recognition system
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
- Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
- General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
- DeepSeek-OCR 2: Visual Causal Flow
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering