Collapse, Not Complexity: Failure-Conditioned Decomposition Repair for End-to-End Document Parsing
cs.CV, cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 5 pages, 3 figures, 2 tables. Submitted to ICASSP 2027
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
- PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
- Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
- Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
- Pre-Inference Routing for Cost-Efficient Document Field Extraction
- Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study
- DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
- LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR
- How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- GLM-OCR Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models