Mind2Report: Expert-Level Commercial Report Synthesis via Cognitive Deep Research Agent
cs.CL
Submitted: 2026-01-08
Updated: 2026-08-26
Comments: Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026), Rome, Italy
Code: https://github.com/Melmaphother/Mind2Report
License: http://creativecommons.org/licenses/by/4.0/
The gist: Synthesizing informative commercial reports from massive and noisy web sources is critical for high-stakes business decisions.
Terminology
Abstract
Synthesizing informative commercial reports from massive and noisy web sources is critical for high-stakes business decisions. Although recent deep research agents (DRAs) achieve notable progress, their reports remain limited in quality, reliability, and coverage. These mainly stem from ambiguous intents that cause search drift, retrieved web content that rapidly saturates the context window, and single-pass synthesis that limits report comprehensiveness. In this work, we propose Mind2Report, a cognitive deep research agent that emulates commercial analysts to synthesize expert-level reports. Mind2Report first probes fine-grained commercial intent to establish a structured outline, then recursively explores web sources and distills validated evidence into research memory to preserve context efficiency. Meanwhile, the research memory and outline continuously co-evolve, refining the report structure to avoid rigid initial planning. Finally, Mind2Report iteratively synthesizes the report based on the evolving outline and accumulated evidence. Together, these designs enable reliable and context-efficient long-horizon commercial deep research. To rigorously evaluate commercial DRAs, we further construct QRC-Eval, comprising 200 real-world commercial tasks and a holistic evaluation framework covering report quality, reliability, and coverage. Extensive experiments demonstrate that Mind2Report consistently outperforms leading proprietary and open-source DRAs, while ablations verify the effectiveness of each module and further analyze the challenges they address. We expect this work to advance the development of commercial deep research agents.
Sources
- GPT-4 Technical Report
- A Survey on Knowledge-Oriented Retrieval-Augmented Generation
- Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
- Tongyi DeepResearch Technical Report
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- DeepSeek-V3 Technical Report
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
- Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
- How Far Are We from Genuinely Useful Deep Research Agents?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering