Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding
cs.CL, cs.CV
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/rednote-hilab/dots.ocr
License: http://creativecommons.org/licenses/by/4.0/
The gist: Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed.
Terminology
Abstract
Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation, OCR outputs are strongly grounded in the input image, making diffusion-based parallel generation promising. However, when several tokens are predicted in one diffusion step, each is predicted before the others are known. Committing them directly can therefore introduce errors. We therefore introduce GravityOCR, a parameter-shared AR-block-diffusion model jointly trained for parallel drafting and causal AR verification. Verifying drafts before commitment lets the model commit multiple output tokens per round without a separate drafting network. The causal AR path also enables GRPO with sequence- and structure-level OCR rewards, avoiding diffusion-trajectory likelihood estimation while updating the shared drafter parameters. On OmniDocBench v1.6, AR-path GRPO improves the Overall score from 94.92 to 95.16 without reducing diffusion drafting efficiency, while the final model remains close to the original GLM-OCR score of 95.48. In an SGLang serving deployment, GravityOCR commits an average of 9.7 output tokens per forward pass and achieves a 3.94 times decode-only speedup on region crops and a 1.32 times end-to-end page-processing speedup over AR decoding.
Sources
- Nougat: Neural Optical Understanding for Academic Documents
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
- Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
- PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
- PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
- MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
- GLM-OCR Technical Report
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
- EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Building and better understanding vision-language models: insights and future directions
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
- HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding
- TiDAR: Think in Diffusion, Talk in Autoregression
- OvisOCR2 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering