Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion
cs.CL, cs.DL
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/grobidOrg/grobid
License: http://creativecommons.org/licenses/by/4.0/
The gist: Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems.
Terminology
Abstract
Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text. GROBID, a modular font-stream parser running on CPU, is the de-facto standard for structuring scientific articles and underpins several of the largest open scholarly corpora. We pair it with a lightweight CPU detector localising figure, table, and paratext (header, footer, page number) regions, encoded as typed-area masks whose tokens are routed to GROBID's specialised models or discarded. On two PMC corpora, Bioinformatics (1,926 articles) and Materials Science (2,595), scored against JATS with a section-aware structural protocol, our extension improves over plain GROBID on most metrics (NS +0.025 / +0.013; +0.086 paragraph recall on Materials Science, d z = 1.08), and caption-linked figure recovery improves on both corpora. On the external Table-BRGM benchmark, table detection recovers F1 0.16 to 0.94 and table structure follows (GriTS-Top 0.27 to 0.78, below the strongest GPU system). On body text, against four vision-based systems (Docling, MinerU, olmOCR, dots.ocr), it has the best paragraph precision on both corpora, the best section detection on Materials Science, and a character error rate within 0.004 of the best GPU parser. End-to-end on CPU, it costs 2.7 -- 3.2 times less than the cheapest GPU system (Docling) and 10 -- 14 times less than generative parsers.
Sources
- GLM-OCR Technical Report
- SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
- Nougat: Neural Optical Understanding for Academic Documents
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
- Efficient Document Parsing via Parallel Token Prediction
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
- olmOCR 2: Unit Test Rewards for Document OCR
- Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
- Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents
- PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
- FireRed-OCR Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering