Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

arXiv:2502.20295 · cs.LG, cs.AI, cs.CV · Submitted 2026-08-07 · Read on arXiv

Benjamin Gutteridge, Matthew Jackson, Toni Kukurin, Xiaowen Dong

University of Oxford · QuantCo

cs.LG, cs.AI, cs.CV

Submitted: 2026-08-07

Updated: 2026-08-10

Comments: 10 pages (36 including references and appendices), 11 figures, accepted at COLM 2026, earlier version accepted at AAAI 2025 Workshop on Document Understanding and Intelligence

Code: https://github.com/BenGutteridge/judge-a-book-by-its-cover

Project page: https://azure.microsoft.com/en-gb/pricing/details/cognitive-services/computer-vision

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 68/100

The gist: This paper, "Judge a Book by Its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription," investigates the effectiveness of combining Optical Character Recognition

Terminology

Summary

This paper, Judge a Book by Its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription, investigates the effectiveness of combining Optical Character Recognition (OCR) with Multi-Modal Large Language Models (MLLMs) for the zero-shot transcription of multi-page handwritten documents. The authors note that while Handwriting text recognition (HTR) remains a challenging task, existing approaches often fail to leverage the shared context of multi-page documents, such as semantic content and handwriting style across pages, because MLLMs are typically used for transcription at the page level, meaning they throw away this shared context.

To address this, the paper proposes two specific prompting strategies:

  • OCR + PAGE 1: This strategy provides the MLLM with the OCR output for the entire document as well as a single page image, where the page image chosen is always the first page of the document.

  • OCR + PAGE N: This strategy provides the chosen page image, and the OCR transcription of every page, where the page image is chosen by an upstream (cheap) LLM to be the most beneficial for transcription.

The authors introduce a benchmark for this task consisting of four datasets: Malvern-Hills, a new multi-page handwritten document dataset derived from previously undigitised archival sources, and three synthesized datasets: IAM, Bentham (comprising 18th-century notes by Jeremy Bentham), and CASIA-5 (a Chinese handwriting dataset). The research evaluates a comprehensive suite of methods including page-by-page (PBP) and all-at-once (AAO) processing, using models such as OpenAI’s GPT-4 O, Google’s GEMINI-2.5- PRO, and the open-source GEMMA-3-27 B.

The experimental results reveal that the best-performing method overall is either OCR + PAGE 1 or OCR + PAGE N. Despite having access to less than half, and sometimes as little as a quarter, of the original source data, the document images, these methods outperform methods that have access to all of the images, and those with access to both the OCR text and all of the images. In terms of semantic accuracy, the authors found that our methods have over 20% fewer major errors (genuine mistakes or hallucinations) than the next best method.

The paper further demonstrates the scalability and generalization of these methods:

  • Scaling: When scaling to longer documents, OCR + PAGE 1 remains the top or joint-top method across all three [long-document variants], even when "having access to <20% (or even <10%) of document images."

  • Generalization: The findings generalize to non-Latin script, as demonstrated on the CASIA-5 dataset where OCR + PAGE X achieves 8.94% CER, nearly matching the more expensive OCR+IMAGES (8.39%) at lower token cost.

The authors conclude that a single page is often as good as the full document and that their proposed methods improve transcription accuracy while balancing cost and performance, effectively leveraging common context from limited, expensive image input to improve correction of cheap text input.

Improvements for AI systems

1. Hybrid Visual-Textual Transcription Architecture

  • The Improvement: Transition from a page-by-page visual processing model to a sparse-visual, dense-textual architecture. This involves decoupling the visual input from the textual input, where the model receives the full text via cheap OCR and uses a single, high-fidelity image as a visual anchor for the entire document.

  • What the improved system can do: It can transcribe massive, multi-page archives (e.g., 100+ page ledgers) with extremely high semantic accuracy and minimal hallucination, while reducing computational costs and token usage by over 90% compared to traditional vision-heavy models.

2. Intelligent Visual Sampling Agent (Upstream Page Selector)

  • The Improvement: Integrate an upstream lightweight LLM agent designed specifically to analyze raw OCR text for uncertainty, layout shifts, or semantic gaps. Instead of feeding the first page to a Multi-Modal LLM (MLLM), this agent identifies the most information-dense or visually representative page.

  • What the improved system can do: It can dynamically select the optimal page to show the MLLM—such as a page with complex character ligatures or a page where the OCR confidence is lowest—ensuring that the single visual input provides the maximum possible corrective context for the rest of the document.

3. Global Handwriting Style Embedding

  • The Improvement: Develop a feature extraction module that uses a single selected page image to generate a global handwriting style embedding (capturing slant, pressure, and character formation). This embedding is then injected into the transformer layers while processing the OCR text of all subsequent pages.

  • What the improved system can do: It can maintain stylistic consistency across a long document, allowing the AI to learn a specific person's handwriting from one page and use that knowledge to correct OCR errors on every other page, even if those pages are never visually processed.

4. Cost-Optimized Multi-Script Digitization Pipeline

  • The Improvement: Implement a low-vision, high-text pipeline for non-Latin and complex character scripts (like Chinese or historical scripts) that prioritizes OCR-text correction over visual character recognition.

  • What the improved system can do: It can perform large-scale digitization of global historical archives at a fraction of the current cost, achieving accuracy levels nearly identical to expensive, vision-intensive models while significantly lowering the barrier for digitizing non-Latin historical documents.

5. Hallucination-Resistant Semantic Error Correction

  • The Improvement: A two-stage refinement process where the first stage performs raw OCR and the second stage uses the OCR + Page X prompting strategy to perform semantic validation.

  • What the improved system can do: It can specifically target and eliminate major errors (genuine mistakes where the model invents text) by using the single page image to cross-reference the semantic flow of the OCR text, reducing critical transcription errors by more than 20%.

Abstract

Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or rely on zero-shot tools such as OCR engines and multi-modal LLMs (MLLMs). MLLMs have shown promise both as end-to-end transcribers and as OCR post-processors, but to date there is little empirical research evaluating different MLLM prompting strategies for HTR, particularly for the case of multi-page documents. Most handwritten documents are multi-page, and share context such as semantic content and handwriting style across pages, yet MLLMs are typically used for transcription at the page level, meaning they throw away this shared context. They are also typically used as either text-only post-processors or image-only OCR alternatives, rather than leveraging multiple modes. This paper investigates a suite of methods combining OCR, LLM post-processing and MLLM end-to-end transcription, for the task of zero-shot multi-page handwritten document transcription. We introduce a benchmark for this task from existing single-page datasets, including a new dataset, Malvern-Hills. Finally, we introduce OCR+PAGE-1 and OCR+PAGE-N, prompting strategies for multi-page transcription that outperform existing methods by sharing content across pages while minimizing prompt complexity.

Sources

Related papers