Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
cs.CL, cs.CV
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 15 pages, 5 figures, 8 tables. Model at https://huggingface.co/jinaai/jina-ocr-v1
Code: https://github.com/rednote-hilab/dots.ocr
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs.
Terminology
Abstract
We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at https://huggingface.co/jinaai/jina-ocr-v1.
Sources
- Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
- Qwen2.5-VL Technical Report
- Nougat: Neural Optical Understanding for Academic Documents
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- DFlash: Block Diffusion for Flash Speculative Decoding
- PaLM: Scaling Language Modeling with Pathways
- PaddleOCR 3.0 Technical Report
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- GLM-OCR Technical Report
- Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- Infinity-Parser2 Technical Report
- Segment Anything
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering