WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
cs.CV
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/Tencent/WeVisDoc
Project page: https://tencent.github.io/WeVisDoc
Terminology
Sources
- Qwen3-VL Technical Report
- PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
- Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
- UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
- GLM-OCR Technical Report
- Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
- STEP3-VL-10B Technical Report
- Infinity-Parser2 Technical Report
- Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
- How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings
- POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
- Ovis2.5 Technical Report
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- olmOCR 2: Unit Test Rewards for Document OCR
- FireRed-OCR Technical Report
- HunyuanOCR Technical Report
- Kimi K2.5: Visual Agentic Intelligence
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models