Dense Process Supervision for Search Agents via Fact Utility Estimation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Dense Process Supervision for Search Agents via Fact Utility Estimation".
Jane: The paper was written by Rongzhi Zhu, Xiangyu Liu, Yi Liu, Shuo Zhang, Ruirui Zhang et al. from Nanjing University and Ant Group and National Institute of Healthcare Data Science.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: So, looking at the title, "Dense Process Supervision for Search Agents via Fact Utility Estimation," what does that tell us right away?
Jane: It suggests we’re moving away from just waiting until the very end to judge success. Instead of just looking at the final answer, we are going to supervise or guide the *process* of finding that answer.
Lu: And that's tied directly into "Fact Utility Estimation," meaning we have a way to measure how helpful each piece of data is before it even contributes to the final score.
Meng: From an engineering standpoint, this addresses the sparse reward problem; we’re not waiting for a massive final reward signal when the agent could be receiving small, useful feedback much earlier in the steps.
Lalam: It really implies that we are acknowledging that intelligence is built by accumulation—by building up a reliable store of facts as we go.
Summary/Methodology: Tom: This paper introduces FactAgent to solve that problem, and it’s pretty clever how they handle the way LLMs interact with their environment.
Jane: They replace the standard, messy interaction history—which is just a long string of text—with a structured "fact store."
Lu: That fact store is key because it allows us to keep track of all the evidence in a way that is compact and organized, which helps manage those long, complex reasoning chains.
Meng: When the agent performs an 'Assert' action, it’s not just dumping raw text; it's distilling the observation into structured triples and putting them into that store.
Lalam: That move from unstructured text to a semantic fact store shows that we are forcing the AI to be more intentional about what data it chooses to retain for a better final outcome.
Improvements/Results: Tom: This is where it gets really interesting, looking at how they assign credit for those 'Assert' and 'Search' steps.
Jane: They use a Bayesian approach to estimate the utility of groups of semantically equivalent facts, which is much more robust than just counting successes.
Lu: The clustering technique allows us to share statistical strength across related facts, so if two different phrasing of the same thing is found, we treat them as one strong piece of evidence.
Meng: This approach generates "dense process rewards," which translates into fine-grained step-level feedback that's directly usable in their GRPO training setup.
Lalam: The results in Table one and Table two show that this mechanism is much better than just relying on the final answer, indicating a massive leap toward achieving true, step-by-step understanding.
Conclusion: Tom: It’s clear that FactAgent's ability to manage evidence is what drives its success across all seven QA benchmarks.
Jane: It seems like this method is proving that the way we structure the internal knowledge of an AI matters just as much as how we train it.
Lu: The fact utility estimation provides a level of internal transparency and creativity in reasoning that I think will unlock so many new possibilities in problem-solving AI.
Meng: For me, this means agents can actually handle more complex tasks because they aren't just "guessing" based on a long history; they are actively building a verified evidence base.
Lalam: We have to realize that this isn't just an academic win; it’s the promise of making AI truly capable of solving the world’s most intricate problems in a way that feels more reliable and understandable.
Rongzhi Zhu, Xiangyu Liu, Yi Liu, Shuo Zhang, Ruirui Zhang, Rui Wu, Tao Jiang, Zequn Sun,, Wenhao Xu,, Wei Hu,
Nanjing University · Ant Group · National Institute of Healthcare Data Science
cs.CL, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Accepted in the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: This paper introduces a novel framework for enhancing search agents by implementing "Dense Process Supervision for Search Agents via Fact Utility Estimation." The work addresses the critical need to
Key concepts
- Dense Process Supervision
- This approach guides and supervises the entire process of finding an answer, rather than waiting for a final result. It generates fine-grained, step-level feedback (dense process rewards) that allows AI agents to receive small, useful guidance early in their reasoning chain.
- Fact Utility Estimation
- This is a method used to measure how helpful each piece of data is. It uses a Bayesian approach and clustering to share statistical strength across related facts, ensuring the AI recognizes and values high-quality information.
- Fact Store
- Instead of relying on unstructured, long strings of raw text history, this system replaces it with an organized 'fact store.' This structure allows the AI to compactly track and manage complex reasoning chains by storing evidence in a semantic format.
Terminology
Summary
This paper introduces a novel framework for enhancing search agents by implementing Dense Process Supervision for Search Agents via Fact Utility Estimation.
The work addresses the critical need to guide complex reasoning agents—which rely on iterative searches and fact assertions—by quantifying the informational value of intermediate steps. By estimating utility, the model learns to prioritize formative evidence facts rather than exponentially growing intermediate states,
thereby significantly improving performance across multiple question-answering benchmarks.
Fact Utility Estimation Methodology
The core contribution involves analyzing the structure of utility estimation by examining fact-cluster distributions across various datasets (NQ, HotpotQA, TriviaQA). The authors analyze utility by clustering asserted facts and normalizing them into probability masses. A key finding is that utility distributions remain highly concentrated rather than uniformly sparse,
suggesting that effective reasoning relies on limited, high-value evidence. Specifically, the analysis shows that The first three clusters already explain more than 93% of total utility mass across all datasets,
empirically supporting the design assumption that successful search trajectories depend on a small number of highly informative evidence facts.
Performance and Sensitivity Analysis
The effectiveness of the proposed method is demonstrated through comprehensive quantitative analyses. The authors report a sensitivity analysis regarding the aggregation weight, finding that while increasing improves performance up to a point, as becomes too large, the dominance of process rewards can negatively affect overall performance, leading to performance degradation.
Based on this observation, they establish = 0.5 as the default setting in all experiments.
Furthermore, the utility distribution analysis groups questions by difficulty; while Hard questions require slightly more evidence aggregation,
the utility distributions remain strongly concentrated across all levels.
Comparison with Related Process Supervision Methods
FactAgent is rigorously compared against representative process-supervision approaches, including E-GRPO (trajectory-level entity supervision) and VinePPO (Monte Carlo rollout expansion). The paper highlights that FactAgent offers distinct advantages: it performs step-level utility estimation without requiring additional annotation
compared to E-GRPO. Crucially, FactAgent also avoids the computational burden of expensive perstep branching and Monte Carlo rollout expansion associated with VinePPO, while simultaneously achieving stronger performance across benchmarks.
Empirical Results on Benchmarks
The quantitative results demonstrate robust performance gains. Comparing FactAgent against competitors in Table 12 shows superior average EM scores across major datasets (NQ, HotpotQA, TriviaQA). For instance, the average EM score for FactAgent is reported as 45.8 across the analyzed benchmarks, demonstrating a clear improvement over related methods like E-GRPO and VinePPO. The analysis of fact clusters further confirms this strength: the cumulative probability mass captured by top-ranked fact clusters shows high values (e.g., 0.973 on NQ for FactAgent), confirming its ability to identify highly informative evidence facts necessary for successful reasoning trajectories.
Improvements for AI systems
1. Implementing Structured Process-Level Reward Fine-Tuning (RLHF/RLAIF)
-
Improvement: Integrate a process-level reward mechanism, parameterized by an aggregation weight, during Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This method explicitly rewards the quality of the reasoning steps (process) rather than relying solely on the final outcome.
-
Mechanism: The training objective must balance two components: L total = L outcome + times L process.
-
L outcome: Standard reward based on final answer correctness (e.g., EM score).
-
L process: A derived reward function that measures the informational utility or logical coherence of intermediate steps (e.g., successful fact assertion, relevance of search queries).
-
System Capability: The improved system can perform complex, multi-step reasoning tasks (like QA or code generation) and significantly improve reliability even when the final ground truth is sparse or difficult to measure. By tuning (e.g., using =0.5 as a default), the system can be optimized to prioritize robust, verifiable internal thought processes over merely guessing the correct final answer, drastically reducing
hallucinated
reasoning steps.
2. Fact-Cluster Utility-Driven Evidence Extraction and Reasoning
-
Improvement: Replace general information retrieval with a utility-driven evidence aggregation mechanism that focuses on identifying and clustering highly informative facts (fact clusters).
-
Mechanism: When processing an observation or retrieved documents, the system must not treat all facts equally. Instead, it should:
-
Cluster asserted triples into
fact clusters.
-
Calculate a utility score for these clusters (e.g., how much of the total informational mass they capture).
-
Prioritize subsequent actions (assertions, search queries) based on maximizing the cumulative cluster mass captured by only the top N most informative clusters (e.g., N=3 explains >93% of utility).
- System Capability: The AI can perform highly focused, efficient research and QA. When given a large corpus of documents, it will automatically disregard redundant or weakly related information and synthesize answers using only the core, high-utility evidence facts, minimizing cognitive load and drastically reducing the risk of mixing irrelevant details into the final answer.
3. Advanced Modular Agent Architecture (FactAgent Paradigm)
-
Improvement: Design a structured agent that explicitly separates reasoning into three modular, sequential steps: Think to Search/Assert to Answer. This architecture must enforce strict internal formatting and information flow control.
-
Mechanism:
-
Mandatory Internal Monologue (
...): The system must first generate a detailed thought process, citing specific observations and facts, before generating any action. This forces self-correction and explainability. -
Structured Action Execution (JSON): All external actions (Search, Assert, Answer) must be executed via a strictly defined JSON schema.
-
Integrated Fact Store: The system must maintain an active
fact store
derived from successfulAssertactions and actively query this store during the finalAnswergeneration.
- System Capability: This creates a highly reliable, auditable, and performant research agent suitable for enterprise applications. It guarantees that every action is traceable back to a thought process and verifiable evidence, making it robust against model drift or unstructured output failures.
4. Dynamic Difficulty-Aware Prompting (Meta-Prompting)
-
Improvement: Implement a meta-prompt layer that dynamically adjusts the agent's operational strategy and required depth of reasoning based on the perceived difficulty of the input query or task dataset.
-
Mechanism: Before execution, the system estimates difficulty (e.g., based on historical success rates or initial search results).
-
Easy/Medium Difficulty: Focus heavily on direct assertion and quick fact retrieval (low emphasis).
-
Hard Difficulty: Mandate iterative refinement, requiring multiple
Searchsteps, explicit evidence aggregation across multiple clusters, and higher process reward weighting to ensure thorough exploration. -
System Capability: The AI transitions from a simple question-answering bot into a sophisticated research assistant. It knows when to stop searching (for easy questions) versus when it must exhaustively explore all logical paths and evidence sources (for hard questions), leading to consistent performance across varying task complexities.
Abstract
Reinforcement learning (RL) for search agents typically relies on outcome rewards. However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. It is hard to separate their contributions from the final result. In this paper, we propose a dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts. We first extract structured facts from raw observations and organize them into an explicit fact store. To support credit assignment, we then cluster semantically equivalent facts and infer the posterior utility of each fact cluster using Bayesian estimation over group rollouts. Finally, we convert the estimated fact utilities into dense step-level rewards to guide RL training. Experiments on seven single-hop and multi-hop QA benchmarks show that our method consistently outperforms existing baselines. Ablation studies validate clear relative improvements on multi-hop QA compared to outcome reward-only training.
Sources
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Group-in-Group Policy Optimization for LLM Agent Training
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- Step-DeepResearch Technical Report
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- VinePPO: Refining Credit Assignment in RL Training of LLMs
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
- DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
- Humanity's Last Exam
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- Tongyi DeepResearch Technical Report
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- WebDancer: Towards Autonomous Information Seeking Agency
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering