Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-09-21
Code: https://github.com/marin-community/marin
License: http://creativecommons.org/licenses/by/4.0/
The gist: PDF corpora advertise their size in tokens, but every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) is computed per document, and none decomposes its token total.
Terminology
Abstract
PDF corpora advertise their size in tokens, but every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) is computed per document, and none decomposes its token total. Because PDF length is extremely skewed, the two units can describe the same corpus very differently. We ask how the headline statistics of a web-PDF corpus change when each document is weighted by the text it contributes rather than counted once. We used CC-MAIN-2021-31-PDF-UNTRUNCATED (7.9M Common Crawl PDFs, 32.6B tokens), the one public corpus that pairs the fragments Common Crawl stored with the re-fetched originals. Text mass is highly concentrated: 3.02% of text-bearing documents hold half the tokens (Gini 0.807). The clearest consequence is Common Crawl's payload cap, which truncated 23.06% of these documents but 63.08% of their text. Reconstructing the truncated fragments and extracting both versions, two widely used text-layer parsers recover only 1.4% and 11.4% of that exposed text, so roughly 55-62% of the corpus's text is unrecoverable from the crawl by such pipelines; under the 5MiB cap adopted in March 2025, 30.19% of tokens would still be exposed. We recommend that corpus statistics be reported in both units, documents and tokens.
Sources
- GovScape: A Public Multimodal Search System for 70 Million Pages of Government PDFs
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering