SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
cs.CV, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/huggingface/accelerate
Terminology
Sources
- ScreenAI: A Vision-Language Model for UI and Infographics Understanding
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- PaliGemma: A versatile 3B VLM for transfer
- Nougat: Neural Optical Understanding for Academic Documents
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- RLHF Workflow: From Reward Modeling to Online RLHF
- PP-OCR: A Practical Ultra Lightweight OCR System
- OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
- LLaVA-OneVision: Easy Visual Task Transfer
- Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models