Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
Benjamin Gutteridge, Matthew Jackson, Toni Kukurin, Xiaowen Dong
University of Oxford · QuantCo
cs.LG, cs.AI, cs.CV
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 10 pages (36 including references and appendices), 11 figures, accepted at COLM 2026, earlier version accepted at AAAI 2025 Workshop on Document Understanding and Intelligence
Code: https://github.com/BenGutteridge/judge-a-book-by-its-cover
Project page: https://azure.microsoft.com/en-gb/pricing/details/cognitive-services/computer-vision
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 68/100
The gist: This paper, "Judge a Book by Its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription," investigates the effectiveness of combining Optical Character Recognition
Terminology
Summary
This paper, Judge a Book by Its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription,
investigates the effectiveness of combining Optical Character Recognition (OCR) with Multi-Modal Large Language Models (MLLMs) for the zero-shot transcription of multi-page handwritten documents. The authors note that while Handwriting text recognition (HTR) remains a challenging task,
existing approaches often fail to leverage the shared context of multi-page documents, such as semantic content and handwriting style across pages,
because MLLMs are typically used for transcription at the page level, meaning they throw away this shared context.
To address this, the paper proposes two specific prompting strategies:
-
OCR + PAGE 1: This strategy provides the MLLM with
the OCR output for the entire document as well as a single page image,
wherethe page image chosen is always the first page of the document.
-
OCR + PAGE N: This strategy provides
the chosen page image, and the OCR transcription of every page,
where the page image ischosen by an upstream (cheap) LLM
to be the most beneficial for transcription.
The authors introduce a benchmark for this task consisting of four datasets: Malvern-Hills, a new multi-page handwritten document dataset derived from previously undigitised archival sources,
and three synthesized datasets: IAM,
Bentham
(comprising 18th-century notes by Jeremy Bentham), and CASIA-5
(a Chinese handwriting dataset). The research evaluates a comprehensive suite of methods
including page-by-page (PBP)
and all-at-once (AAO)
processing, using models such as OpenAI’s GPT-4 O,
Google’s GEMINI-2.5- PRO,
and the open-source GEMMA-3-27 B.
The experimental results reveal that the best-performing method overall is either OCR + PAGE 1 or OCR + PAGE N.
Despite having access to less than half, and sometimes as little as a quarter, of the original source data, the document images,
these methods outperform methods that have access to all of the images, and those with access to both the OCR text and all of the images.
In terms of semantic accuracy, the authors found that our methods have over 20% fewer major errors (genuine mistakes or hallucinations) than the next best method.
The paper further demonstrates the scalability and generalization of these methods:
-
Scaling: When
scaling to longer documents,
OCR + PAGE 1 remains the top or joint-top method across all three [long-document variants],
even when "having access to <20% (or even <10%) of document images." -
Generalization: The findings
generalize to non-Latin script,
as demonstrated on the CASIA-5 dataset whereOCR + PAGE X achieves 8.94% CER, nearly matching the more expensive OCR+IMAGES (8.39%) at lower token cost.
The authors conclude that a single page is often as good as the full document
and that their proposed methods improve transcription accuracy while balancing cost and performance,
effectively leveraging common context from limited, expensive image input to improve correction of cheap text input.
Improvements for AI systems
1. Hybrid Visual-Textual Transcription Architecture
-
The Improvement: Transition from a
page-by-page
visual processing model to asparse-visual, dense-textual
architecture. This involves decoupling the visual input from the textual input, where the model receives the full text via cheap OCR and uses a single, high-fidelity image as avisual anchor
for the entire document. -
What the improved system can do: It can transcribe massive, multi-page archives (e.g., 100+ page ledgers) with extremely high semantic accuracy and minimal hallucination, while reducing computational costs and token usage by over 90% compared to traditional vision-heavy models.
2. Intelligent Visual Sampling Agent (Upstream Page Selector)
-
The Improvement: Integrate an
upstream
lightweight LLM agent designed specifically to analyze raw OCR text for uncertainty, layout shifts, or semantic gaps. Instead of feeding the first page to a Multi-Modal LLM (MLLM), this agent identifies the mostinformation-dense
orvisually representative
page. -
What the improved system can do: It can dynamically select the optimal page to
show
the MLLM—such as a page with complex character ligatures or a page where the OCR confidence is lowest—ensuring that the single visual input provides the maximum possible corrective context for the rest of the document.
3. Global Handwriting Style Embedding
-
The Improvement: Develop a feature extraction module that uses a single selected page image to generate a
global handwriting style embedding
(capturing slant, pressure, and character formation). This embedding is then injected into the transformer layers while processing the OCR text of all subsequent pages. -
What the improved system can do: It can maintain stylistic consistency across a long document, allowing the AI to
learn
a specific person's handwriting from one page and use that knowledge to correct OCR errors on every other page, even if those pages are never visually processed.
4. Cost-Optimized Multi-Script Digitization Pipeline
-
The Improvement: Implement a
low-vision, high-text
pipeline for non-Latin and complex character scripts (like Chinese or historical scripts) that prioritizes OCR-text correction over visual character recognition. -
What the improved system can do: It can perform large-scale digitization of global historical archives at a fraction of the current cost, achieving accuracy levels nearly identical to expensive, vision-intensive models while significantly lowering the barrier for digitizing non-Latin historical documents.
5. Hallucination-Resistant Semantic Error Correction
-
The Improvement: A two-stage refinement process where the first stage performs raw OCR and the second stage uses the
OCR + Page X
prompting strategy to perform semantic validation. -
What the improved system can do: It can specifically target and eliminate
major errors
(genuine mistakes where the model invents text) by using the single page image to cross-reference the semantic flow of the OCR text, reducing critical transcription errors by more than 20%.
Abstract
Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or rely on zero-shot tools such as OCR engines and multi-modal LLMs (MLLMs). MLLMs have shown promise both as end-to-end transcribers and as OCR post-processors, but to date there is little empirical research evaluating different MLLM prompting strategies for HTR, particularly for the case of multi-page documents. Most handwritten documents are multi-page, and share context such as semantic content and handwriting style across pages, yet MLLMs are typically used for transcription at the page level, meaning they throw away this shared context. They are also typically used as either text-only post-processors or image-only OCR alternatives, rather than leveraging multiple modes. This paper investigates a suite of methods combining OCR, LLM post-processing and MLLM end-to-end transcription, for the task of zero-shot multi-page handwritten document transcription. We introduce a benchmark for this task from existing single-page datasets, including a new dataset, Malvern-Hills. Finally, we introduce OCR+PAGE-1 and OCR+PAGE-N, prompting strategies for multi-page transcription that outperform existing methods by sharing content across pages while minimizing prompt complexity.
Sources
- GPT-4 Technical Report
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Kleister: A novel task for Information Extraction involving Long Documents with Complex Layout
- The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
- OCR-free Document Understanding Transformer
- FABLES: Evaluating faithfulness and content selection in book-length summarization
- A Comprehensive Survey on Long Context Language Modeling
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- ANLS* -- A Universal Document Processing Metric for Generative Large Language Models
- Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
- Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study
- A Survey of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks