Prefix-Adaptive Block Diffusion for Efficient Document Recognition
cs.CV, cs.AI
Submitted: 2026-05-16
Updated: 2026-08-29
Comments: 16pages,6 figures
Code: https://github.com/SII-sc22mc/PA-BDM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing.
Terminology
Abstract
Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind denoising and cache commitment to fixed block boundaries: parallelism shrinks during intra-block denoising, while generated tokens cannot be cached until the whole block is completed. Moreover, intra-block bidirectional denoising conflicts with inter-block autoregression, creating inconsistent information flow that can challenge structure-sensitive recognition. We propose the Prefix-Adaptive Block Diffusion Model (PA-BDM), which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit. PA-BDM uses Confidence-gated Structural Loss (CSL) to build low-entropy prefixes before extending training to longer continuations. During inference, Progressive Prefix Commitment (PPC) then dynamically commits the longest reliable prefix into the KV cache and resets the next candidate range from the updated prefix, restoring a large parallel decoding space at each step. Experiments show that the 3B PA-BDM achieves higher recognition scores on several benchmarks and improves inference throughput by 71.6% over the 2.5B MinerU-Diffusion.
Sources
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Qwen2.5-VL Technical Report
- Nougat: Neural Optical Understanding for Academic Documents
- PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
- PaddleOCR 3.0 Technical Report
- MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
- MDiff4STR: Mask Diffusion Model for Scene Text Recognition
- GLM-OCR Technical Report
- Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- Fast Inference from Transformers via Speculative Decoding
- LaViDa: A Large Diffusion Language Model for Multimodal Understanding
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
- DODO: Discrete OCR Diffusion Models
- Large Language Diffusion Models
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
- FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models