DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/VXRealLimited/DocMIDE
License: http://creativecommons.org/licenses/by/4.0/
The gist: Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page.
Terminology
Abstract
Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values requires multi-hop derivation, such as aggregating sub-categories or reasoning over visual marks. While existing methods handle explicit text spans or simple implicit queries, they fail at multi-hop visual reasoning even after standard fine-tuning: models retrieve incorrect visual evidence, or retrieve it correctly and then skip the intermediate steps of the derivation. To address this, we introduce DocMIDE, a fine-tuning framework that trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. DocMIDE constrains generation to a plan-retrieve-derive structure and optimizes it with Group Relative Policy Optimization under a four-component, rule-based reward that scores output format, the retrieved evidence block, every intermediate derivation step, and the final value against a verified reference trace. On a 4,151-pair implicit extraction benchmark, DocMIDE raises accuracy from 70.8% to 95.9% on Qwen3.5-4B from only a small set of annotated examples, and transfers to a second backbone architecture. Supervised demonstrations alone do not close this gap at any budget we tested; rewarding the intermediate steps is what does.
Sources
- Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
- GLM-OCR Technical Report
- FireRed-OCR Technical Report
- DeepSeek-OCR 2: Visual Causal Flow
- PaddleOCR 3.0 Technical Report
- Unlimited OCR Works
- DocILE Benchmark for Document Information Localization and Extraction
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- LoRA: Low-Rank Adaptation of Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection