InSight-doc: Agentic Visual Perception for Long-Document Understanding
Kaican Li, Weiyan Xie, Lewei Yao, Jiannan Wu, Lanqing Hong, Yongxiang Huang, Nevin L. Zhang
The Hong Kong University of Science and Technology · Huawei
cs.CV, cs.CL, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/m-Just/InSight-doc
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: InSight-doc: Agentic Visual Perception for Long-Document Understanding proposes a novel framework to address the challenges of long-document understanding, which often requires reasoning over many
Terminology
Summary
InSight-doc: Agentic Visual Perception for Long-Document Understanding proposes a novel framework to address the challenges of long-document understanding, which often requires reasoning over many visually rich pages, making inference costly and prone to context rot. The paper introduces InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource. Instead of processing documents at a fixed high resolution, InSight-doc starts from a low-resolution overview and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever.
The core methodology involves training an agent to perform an interleaved multimodal chain-of-thought. As illustrated in Figure 2, the model begins with a low-resolution view of the full document and, in each round, emits a thought and a zoom-in tool call specifying the image index, a region label, and a bounding box. The cropped high-resolution region is then appended as visual evidence, allowing the model to iteratively refine its focus until it can produce a final answer. This coarse-to-fine workflow mimics human reading behavior, starting from a high-level overview and only jumping into details when necessary, thereby reducing context pressure and allowing the model to focus its attention on relevant information.
To train this agent, the authors constructed a large-scale, high-quality training corpus. This corpus includes 17.9K high-quality SFT examples with region-level zoom-in trajectories and 19.2K hard RL examples. The data construction pipeline, detailed in Section 4, involves a three-stage filtering process and a two-agent trajectory generator based on InSight-o3. The filtering stages discard questions answerable without the document, remove questions already answerable at low DPI without zoom, and then construct zoom-in CoTs for the remaining items. The resulting SFT data contains 17,913 trajectories (14,216 answerable and 3,697 unanswerable), and the RL split contains 19,236 prompts. The average document length is 18.51 pages, and SFT trajectories contain an average of 2.61 assistant CoT/tool-use rounds.
The experimental results demonstrate significant improvements across multiple benchmarks. On document VQA benchmarks, InSight-doc-8B (SFT+RL) achieves an average accuracy of 66.9% at low initial resolution (r=0.25), outperforming its base model Qwen3-VL-8B by 16.4 points. At medium resolution (r=0.5), it reaches 72.6% average accuracy, a 4.3-point improvement. The gains are consistent across DUDE, MP-DocVQA, MMLongBench-Doc, and LongDocURL. The paper states: At the low initial resolution (r = 0.25), InSight-doc-8B (SFT+RL) achieves an average accuracy of 66.9%, significantly outperforming its base model, Qwen3-VL-8B, by 16.4 points.
Furthermore, RL contributes substantially beyond SFT alone, increasing average accuracy from 56.6% to 66.9% at r=0.25 and from 64.6% to 72.6% at r=0.5.
On long documents, InSight-doc reduces hallucination by more than 40% and inference latency by 41%–68% while maintaining an accuracy lead. Specifically, on unanswerable questions, InSight-doc achieves F1 scores of 69.1 on DUDE and 74.4 on MMLongBench-Doc at r=0.25, improving over the baseline by 24.6 and 25.9 points, respectively. The paper notes: At r = 0.25, InSight-doc achieves 69.1 on DUDE and 74.4 on MMLongBench-Doc, improving over Qwen3-VL-8B without zoom by 24.6 and 25.9 points, respectively.
This indicates the model is better able to recognize when the document does not contain sufficient evidence.
The efficiency gains are also substantial. InSight-doc achieves comparable or higher accuracy with substantially shorter sequences. At 50 DPI (r=0.25), it obtains 66.9% average accuracy, nearly matching the 68.3% of the 100-DPI baseline while using 58% fewer tokens. At 70 DPI (r=0.35), it reaches 70.6%, exceeding the 69.4% of the 140-DPI baseline while reducing token count by 66%. Latency is similarly reduced: at 70 DPI, InSight-doc outperforms the 140-DPI baseline by 1.2 points on average while reducing latency by 54%. On the longest-document subset, it requires only 11.2 seconds per example compared to 39.3 seconds for the 140-DPI baseline, a 71% reduction while improving accuracy by 3.0 points.
The trajectory quality analysis shows that RL training further improves localization quality and trajectory stability. At r=0.25, RL raises evidence-box coverage to 82.3% while using fewer crops than the base model, and reduces redundant and stuck trajectories from 14.1% and 9.7% to 5.8% and 0.1%, respectively. The paper concludes: "We introduced InSight-doc, which treats visual resolution as an adaptive reasoning-time resource for long-document understanding. By selectively zooming from low-resolution pages into relevant regions, it improves accuracy while reducing, hallucination, context length and end-to-end inference latency." The code, datasets, and model are released at https://github.com/m-Just/InSight-doc.
Improvements for AI systems
Improvements to AI systems:
-
Adaptive visual resolution allocation: Replace fixed high-resolution document processing with a coarse-to-fine, agentic zoom-in mechanism. The system starts at low resolution and selectively crops high-resolution regions only when needed, reducing token usage by 58–66% while maintaining or exceeding accuracy.
-
Interleaved multimodal chain-of-thought with tool calls: Train the model to emit explicit reasoning steps (
thoughts
) interleaved with zoom-in actions (image index, region label, bounding box) in each round. This enables transparent, step-by-step evidence gathering that mimics human reading, improving answer reliability and reducing context rot. -
Hallucination suppression via explicit unanswerability detection: Use the agentic zoom process to force the model to verify evidence before answering. This reduces hallucination on unanswerable questions by over 40%, achieving F1 scores of 69.1 (DUDE) and 74.4 (MMLongBench-Doc) at low resolution—improvements of 24.6 and 25.9 points over baselines.
-
Reinforcement learning for trajectory stability and localization: Apply RL (beyond SFT) to optimize zoom-in sequences, increasing evidence-box coverage to 82.3% while reducing redundant and stuck trajectories from 14.1% and 9.7% to 5.8% and 0.1%, respectively. This yields a 10.3-point accuracy gain at low resolution (56.6% → 66.9%) and 8.0 points at medium resolution (64.6% → 72.6%).
-
Cost-efficient long-document inference: Enable the system to process documents with 18.5 pages on average using only 2.61 tool-use rounds per query. This cuts end-to-end latency by 41–71% (e.g., 39.3s → 11.2s on the longest subset) while improving accuracy by 3.0 points over fixed high-resolution baselines.
-
Data-efficient training pipeline: Use a three-stage filtering process (remove questions answerable without the document, remove those answerable at low DPI, then construct zoom-in trajectories) to generate 17.9K high-quality SFT examples and 19.2K hard RL prompts. This enables training a compact 8B model that outperforms larger fixed-resolution models.
What the improved AI system can do:
-
Answer questions over multi-page, visually rich documents (e.g., scanned reports, charts, tables) with higher accuracy than fixed-resolution models, while using far fewer tokens and less compute.
-
Automatically decide when to zoom in on specific page regions, bounding boxes, or visual elements, without any external retriever or pre-indexing.
-
Explicitly refuse or flag unanswerable questions by verifying that no relevant evidence exists in the document, reducing false positives and improving trustworthiness.
-
Operate in real-time on long documents (e.g., 20+ pages) with latency under 12 seconds per query, making it suitable for interactive document assistants, legal review, and academic Q&A.
-
Scale to even longer documents by keeping initial resolution low and only allocating high-resolution processing to the most relevant regions, avoiding context window overflow.
-
Provide interpretable reasoning traces (thoughts + zoom actions) that show exactly which page regions were examined, enabling auditability and debugging of the model's evidence-gathering process.
Abstract
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource. InSight-doc starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever. To train such an agent, we construct an active-perception corpus of 17.9K high-quality SFT examples with region-level zoom-in trajectories, accompanied by 19.2K hard RL examples. Through SFT+RL, InSight-doc-8B improves the baseline by 4.3--16.4 accuracy points over document VQA benchmarks. On long documents, it reduces hallucination by more than 40% and inference latency by 41%--68% while maintaining an accuracy lead. Our code, datasets, and model are released at https://github.com/m-Just/InSight-doc.
Sources
- ColPali: Efficient Document Retrieval with Vision Language Models
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- Qwen3-VL Technical Report
- Seed1.5-VL Technical Report
- Qwen2.5-VL Technical Report
- SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
- Kimi-VL Technical Report
- GRIT: Teaching MLLMs to Think with Images
- Evaluating Object Hallucination in Large Vision-Language Models
- ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
- HybridFlow: A Flexible and Efficient RLHF Framework
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models