Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
cs.AI, cs.CV
Submitted: 2025-05-24
Updated: 2026-08-28
Code: https://github.com/Doc-CoB/Doc-CoB
License: http://creativecommons.org/licenses/by/4.0/
The gist: Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rely on only a few
Terminology
Abstract
Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rely on only a few relevant layout regions. However, existing methods either adopt a one-pass strategy that implicitly assumes all layouts are equally important, or focus excessively on small regions at the cost of losing critical layout information. To address these limitations, we introduce Doc-CoB (Chain-of-Boxes), a simple-yet-effective framework that integrates coarse-to-fine layout-aware visual reasoning into multimodal large language models. Instead of directly zooming into small regions, Doc-CoB progressively focuses on query-relevant layouts while preserving global document information. Specifically, it first selects key layout boxes and then focuses on them for further understanding with visual prompting. To support this paradigm, we introduce two reasoning tasks for box recognition and box reasoning, with an automatic pipeline that constructs 249k training samples with intermediate visual supervision. Experiments on seven benchmarks with four popular models show that Doc-CoB significantly improves performance, demonstrating its effectiveness and wide applicability. The code and the data are available at https://github.com/Doc-CoB/Doc-CoB.
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- Scaling Spatial Intelligence with Multimodal Foundation Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Document AI: Benchmarks, Models and Applications
- A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
- Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- OpenAI o1 System Card
- Object-level Visual Prompts for Compositional Image Generation
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
- Hierarchical multimodal transformers for Multi-Page DocVQA
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- Qwen2.5 Technical Report
- mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection