When Is Enough Not Enough? Illusory Completion in Search Agents

arXiv:2602.07549 · cs.AI, cs.CL · Submitted 2026-02-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "When Is Enough Not Enough? Illusory Completion in Search Agents".

Jane: Recent search agents often fail to reliably reason across all requirements in multi-constraint problems because they frequently engage in illusory completion,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re diving into this paper today, "When Is Enough Not Enough? Illusory Completion in Search Agents." It’s about how these advanced search agents sometimes get the wrong idea about whether they've actually finished their job when there are multiple conditions to meet.

Jane: Exactly. The core issue here is that these agents can believe a task is done even if some important rules or constraints haven't been properly checked or might have been violated along the way. It’s a problem where the answer looks good on the surface but misses something crucial under closer inspection, which we call underverified answers.

Lu: The authors set up this study by looking specifically at multi-constraint problems, where you have several conditions that all need to be true for an answer to be correct at once. They found that illusory completion happens frequently in these scenarios because the agents struggle to keep track of everything they’ve verified versus what is still left open.

Meng: It sounds like a common hurdle in complex reasoning systems, but what makes this paper interesting is how they are trying to actually see *why* the agent gets it wrong, not just that it got it wrong. They introduce this new tool called the EPISTEMIC LEDGER to track both what evidence is actually available and what the agent thinks about each constraint at every single step of its reasoning process.

Lalam: That sounds like a really detailed way to look inside an agent's head during a search, trying to separate actual support from just hopeful belief, which is where these errors usually hide.

Tom: Right, and that ledger lets them diagnose four specific ways this illusory completion happens: bare assertions, overlooking refutations, stagnation where the agent gets stuck doing the same thing over and over without finding new info on things it hasn't checked yet, and premature exit where it just quits before fully checking everything.

Jane: Those are some pretty concrete failures to track, moving beyond just saying an answer is wrong to figuring out the specific mistake in their reasoning process. It shows that the problem isn't just a final answer error; it’s a structural flaw in how they manage the verification steps throughout their search.

Lu: The paper points out that this failure pattern suggests that these agents lack a structured way of knowing what has been verified and what hasn't across all the candidates they are looking at. It seems like a problem with awareness, not just computation power.

Meng: If an agent isn't tracking its verification state well, it’s going to keep making those kinds of mistakes, so figuring out how to give it better awareness is the next logical step for practical application.

Title and authors: Tom: So, how do they fix this? The authors look at explicitly tracking the constraint state during execution using something called LIVELEDGER and show that this intervention actually makes a big difference in performance.

Jane: And what they found is that explicit constraint-state tracking consistently improves accuracy by up to eleven point six percent and reduces underverified answers by as much as twenty-six point five percent on these multi-constraint problems, which is a substantial gain for reliability.

Lu: LIVELEDGER works by updating the candidate constraint state incrementally along the agent's reasoning path, using both evidential support and the agent's own beliefs to build this ledger. It’s an inference-time tracker that exposes what's happening in real time during execution.

Meng: That sounds like it forces the search agent to be more disciplined about what it accepts as true evidence at each turn, which is exactly what we need for building dependable AI tools for users.

Lalam: It seems like this moves the agent from guessing whether it’s done to actually knowing, step by step, if the necessary conditions are being met or if they’re being ignored.

Tom: The results show that LIVELEDGER has a really strong effect when paired with larger models; for instance, using a 120B backbone led to substantial reductions across those four failure patterns.

Jane: That synergy suggests that the size of the underlying model matters in how effectively it can handle this kind of explicit constraint checking, which is an important detail for anyone trying to deploy these systems.

Lu: They also quantified this behavior using something called the Extent of Candidate Exploration, or ECE, which measures how many different candidates are being explored per reasoning turn. This helps focus the agent on things that genuinely look like they could be the final answer instead of just wandering aimlessly.

Meng: Focusing exploration sounds smart because it reduces wasted computation; if you’re only looking at a small set of promising options, you spend less time chasing dead ends.

Tom: The paper also found that LIVELEDGER makes search more efficient by reducing the number of turns needed for agents like TongyiDR-L-20B, which points toward smoother convergence to a fully verified correct answer.

Jane: So what this means for us listening right now is that when we use these complex AI tools, we can expect them to be more thorough and less likely to give us an answer that looks convincing but isn't actually solid.

Lu: They also provided some interesting qualitative examples showing how the incorrect agents might falsely mark constraints as satisfied while LIVELEDGER correctly flags those violations during the reasoning process.

Meng: That qualitative evidence is super helpful because it shows exactly what that internal tracking looks like when things go wrong versus when they go right.

Title and authors: Lalam: It really grounds the abstract idea of constraint tracking by showing you the actual flow of information as the agent thinks through its steps on a problem.

Tom: So, to wrap up this part of the discussion, the authors are basically saying that explicit constraint checking during execution is a way to stop illusory completion and steer agents away from bad paths toward fully verified answers.

Jane: They argue that this explicit state information helps agents converge more smoothly toward answers that pass every single test, which is vital when you’re dealing with multiple requirements.

Lu: It points toward a future where search agents don't just guess the right direction but actively manage their epistemic state, knowing precisely what they have verified and what they still need to check.

Meng: From an engineering standpoint, it means we can build systems where we know exactly when to stop searching because we’ve hit a verified solution, rather than letting them wander forever on something that's already dead.

Tom: So that’s the big picture for this study on "When Is Enough Not Enough? Illusory Completion in Search Agents." It shows us how tracking those belief and evidence dimensions can fix a fundamental flaw in how search agents handle complex tasks.

Jane: It’s a solid piece of work because it doesn't just point out that agents are sometimes wrong; it gives us the mechanism to see where they fail so we can build better ones.

Lu: This EPISTEMIC LEDGER framework, combined with LIVELEDGER, is a useful way to collect high-quality agent trajectories for future training improvements. It’s like having a detailed diagnostic log for how an AI solves problems.

Meng: Collecting those logs will be key because it gives the researchers the data needed to train models that are inherently better at this kind of verification, not just models that get slightly better at guessing right.

Lalam: I think if we can teach the AI to manage its internal ledger better, it could really help improve how we interact with these powerful reasoning systems in general.

Tom: That’s our time for today on this paper. We’ve looked at the problem of illusory completion and how LIVELEDGER helps us diagnose and mitigate those issues in search agents.

Jane: It really shows that reliability isn't just about getting a good final answer; it’s about having a transparent process behind that answer.

Lu: And this work gives us a new lens for thinking about how to structure reasoning so agents are more aware of the verification landscape.

Meng: We’ll keep watching how these explicit tracking methods start showing up in the next set of practical, real-world AI tools we use every day.

The paper's summary: Tom: So, to wrap up that last bit, this paper is really digging into why search agents sometimes get stuck or give answers that look okay but aren't actually solid when they have to check a bunch of different conditions at once.

Jane: Exactly. They're calling it illusory completion. It means the AI thinks it’s done solving the problem even though some important rules haven't been properly checked or might have been violated along the way, which leads to those underverified answers we talked about before.

Tom: And they show this isn't just a final answer mistake; it's a structural flaw in how these agents manage their verification steps during the whole search process.

Jane: They use this framework called the EPISTEMIC LEDGER to track both what actual evidence is available and what the agent thinks about each constraint at every single step of its reasoning. It’s like giving the AI a detailed log of its own thoughts and evidence as it works through a complex task.

Tom: That ledger lets them diagnose four specific ways this illusory completion happens: bare assertions, overlooking refutations, stagnation where the agent gets stuck doing the same thing over and over without finding new info on things it hasn't checked yet, and premature exit where it just quits before fully checking everything.

Jane: Those are some pretty concrete failures to track that go beyond just saying a final answer is wrong; they reveal the specific mistake in their reasoning process. It suggests the problem isn't just computation power; it’s a lack of awareness about what has been verified and what remains unverified across all the candidates it's looking at.

Tom: If an agent isn't tracking its verification state well, it’s going to keep making those kinds of mistakes, so figuring out how to give it better awareness is the next logical step for practical application.

Jane: The paper points out that this failure pattern suggests these agents lack a structured way of knowing what has been verified and what hasn't across all the candidates they are considering. It seems like a problem with awareness rather than just computational power.

Tom: If an agent isn't tracking its verification state well, it’s going to keep making those kinds of mistakes, so figuring out how to give it better awareness is the next logical step for practical application.

Jane: They look at this because if agents can't manage that internal ledger properly, they just keep guessing and getting stuck in loops or quitting prematurely before they've actually checked everything required.

Tom: And what they suggest to fix this is using something called LIVELEDGER, which is an inference-time tracker that updates the constraint state in real time during execution by combining evidence with the agent’s beliefs.

Jane: That intervention is shown to make a big difference, improving accuracy by up to eleven point six percent and cutting those underverified answers down by as much as twenty-six point five percent on these multi-constraint problems.

The paper's summary: Tom: LIVELEDGER works because it forces the search agent to be more disciplined about what it accepts as true evidence at each turn, steering it away from bad paths toward fully verified answers.

Jane: It shows that explicit constraint-state tracking helps agents converge more smoothly toward answers that pass every single test, which is vital when you're dealing with multiple requirements all at once.

Tom: And they found that this method works really well when paired with larger models; for instance, using a bigger backbone model made the reductions across those four failure patterns even more substantial.

Jane: That synergy suggests that the size of the underlying model matters in how effectively it can handle this kind of explicit checking, which is an important detail for anyone trying to deploy these kinds of systems in real applications.

Tom: They also look at how many different candidates are being explored per turn, using something called ECE, and they found that focusing exploration helps the agent stay on track toward the final answer instead of wandering aimlessly.

Jane: It’s about making sure the AI spends its time looking at things that actually have potential to be the solution, which means less wasted effort.

Tom: The paper also shows how this works qualitatively, with examples where a wrong agent falsely marks constraints as satisfied while LIVELEDGER correctly flags those violations during the reasoning process.

Jane: It’s about showing you exactly what that internal tracking looks like when things go wrong versus when they go right, which makes the whole concept much clearer.

Tom: So in short, this paper argues that explicit constraint checking during execution is a way to stop illusory completion and force agents to be more thorough in their verification steps.

Jane: It really shows that reliability isn't just about getting a good final answer; it’s about having a transparent process behind that answer, which is crucial when you're dealing with complex requirements.

Tom: And this work gives us the mechanism to see where they fail so we can build better systems for these kinds of difficult reasoning tasks.

Jane: This EPISTEMIC LEDGER framework and the LIVELEDGER intervention are tools that help us understand how to structure reasoning so agents are more aware of the verification landscape.

Tom: It points toward a future where search agents don't just guess the right direction but actively manage their internal state, knowing precisely what they have checked and what they still need to check.

Jane: From an engineering standpoint, it means we can build systems where we know exactly when to stop searching because we’ve hit a verified solution, rather than letting them wander forever on something that's already dead.

Tom: So that’s the big picture for this study on illusory completion and how explicit tracking helps fix a fundamental flaw in search agents.

The paper's improvements: Tom: So, after showing us how bad illusory completion can be, the paper isn't just pointing out problems; they’re actually proposing solutions for how we build these agents to prevent them from getting stuck in those loops.

Jane: Exactly. They introduce this concept of LIVELEDGER as an inference-time tracker that updates the constraint state in real time during execution, which is a big shift from just checking things at the very end.

Tom: That’s right, it exposes what's happening with constraint satisfaction right then and there, letting the agent adjust its next move based on what it just found.

Jane: It seems like this mechanism forces the AI to be more disciplined about what evidence it accepts at each step, which is exactly what we need for building more dependable AI tools.

Tom: And they show that this explicit state tracking doesn't just help with accuracy; it makes the search process itself much faster, reducing the total number of turns needed for agents like TongyiDR-L-20B.

Jane: That efficiency gain is interesting because it suggests we can get to a fully verified correct answer much quicker instead of letting the search wander around for ages.

Tom: They also found that this method has a strong effect when you use larger models, meaning bigger models are better equipped to handle this kind of explicit checking effectively.

Jane: The paper flags one limitation though, and they say that while LIVELEDGER is great at tracking what's *found*, it doesn't necessarily fix the problem if the underlying model itself can’t handle complex constraint checks well in the first place.

Tom: That makes sense; you can’t fix a bad foundation with just a better tracker on top of it, right?

Jane: They also mentioned that while LIVELEDGER helps reduce failures like premature exit, there's still room to improve how agents digest all those ledger updates and use that real-time information effectively in their next thought process.

Tom: So the path forward isn't just adding a tracker; it’s about making sure the agent is good at actually using that real-time data to steer its search correctly.

Jane: It moves the focus from just collecting data on failures to improving how we teach agents to manage that internal state better during their thinking process.

Tom: This work suggests that future research needs to focus on making the constraint checking and ledger updates themselves more robust so the agent can truly leverage that information.

Conclusion: Tom: So we’ve covered how illusory completion can mess up search agents when they have to satisfy several constraints at once, and that’s what this paper, "When Is Enough Not Enough? Illusory Completion in Search Agents," is all about.

Jane: It boils down to the idea that these agents sometimes think they're done solving a problem even though some rules haven't been properly checked or might have been violated along the way. It’s this issue of underverified answers popping up even when the final result looks right on paper.

Tom: The authors show us that this isn't just about a wrong final answer; it’s a fundamental flaw in how agents manage their verification steps during the whole search process, which they call epistemic failure.

Jane: They introduce LIVELEDGER to help diagnose this by tracking both what evidence is actually available and what the agent thinks about each constraint at every single step of its reasoning. It’s like giving the AI a detailed log of its own thoughts and evidence as it works through a complex task.

Tom: That ledger lets them pinpoint four specific ways this goes wrong: bare assertions, overlooking refutations, getting stuck in stagnation, and premature exit without checking everything.

Jane: Those are concrete failures to track that move beyond just saying an answer is wrong; they reveal the specific mistake in their reasoning process. It suggests the problem isn't just raw computation; it’s a lack of awareness about what has been verified and what remains unverified across all the candidates it's considering.

Tom: If an agent isn't tracking its verification state well, it’s going to keep making those kinds of mistakes, so figuring out how to give it better awareness is the next logical step for building more dependable AI tools.

Jane: They look at this because if agents can't manage that internal ledger properly, they just keep guessing and getting stuck in loops or quitting prematurely before they've actually checked everything required.

Tom: And what they suggest to fix this is using LIVELEDGER, which is an inference-time tracker that updates the constraint state in real time during execution by combining evidence with the agent’s beliefs.

Jane: That intervention is shown to make a big difference, improving accuracy by up to eleven point six percent and cutting those underverified answers down by as much as twenty-six point five percent on these multi-constraint problems.

Tom: It's interesting that they found this works really well when you use larger models; bigger models are better equipped to handle this kind of explicit checking effectively.

Jane: The paper flags one limitation though, and they say that while LIVELEDGER is great at tracking what's *found*, it doesn't necessarily fix the problem if the underlying model itself can’t handle complex constraint checks well in the first place.

Tom: That makes sense; you can’t fix a bad foundation with just a better tracker on top of it, right?

Jane: They also mentioned that while LIVELEDGER helps reduce failures like premature exit, there's still room to improve how agents digest all those ledger updates and use that real-time information effectively in their next thought process.

Tom: So the path forward isn't just adding a tracker; it’s about making sure the agent is good at actually using that real-time data to steer its search correctly.

Jane: It moves the focus from just collecting data on failures to improving how we teach agents to manage that internal state better during their thinking process.

Tom: This work suggests that future research needs to focus on making the constraint checking and ledger updates themselves more robust so the agent can truly leverage that information.

Lu: From my side, this EPISTEMIC LEDGER is wild because it opens up possibilities for how we model complex, multi-faceted knowledge in AI systems; imagine agents having a true memory of their verification status.

Meng: I’m thinking practically about deployment; if we can build systems that explicitly track these constraints, the reliability for critical applications like autonomous driving or medical diagnosis goes way up.

Lalam: I see this as a step toward more reliable AI interactions overall; when an AI knows what it has verified, it can be much more trustworthy in its outputs and help shape a better digital culture.

Dayoon Ko, Jihyuk Kim, Sohyeon Kim, Haeju Park, Dahyun Lee, Gunhee Kim, Moontae Lee

Seoul National University

cs.AI, cs.CL

Submitted: 2026-02-07

Updated: 2026-10-05

Code: https://github.com/dayoon-ko/illusory_

Importance score: 90/100

The gist: Recent search agents often fail to reliably reason across all requirements in multi-constraint problems because they frequently engage in illusory completion, where they believe tasks are complete

Key concepts

Illusory Completion
This is when an AI agent incorrectly assumes a complex task is finished even though certain necessary conditions or constraints have not been met or violated. It leads to answers that are not fully verified because the agent stops searching prematurely, believing it has found the solution.
EPISTEMIC LEDGER
This is an evaluation framework designed to monitor how an agent reasons across multiple steps. It tracks two things for every constraint: whether the search results actually support it (evidential support) and whether the agent believes it is satisfied (agent belief).
LIVELEDGER
This is a simple intervention that acts as an inference-time tracker. It updates the ledger only based on observation, showing real-time constraint satisfaction status during reasoning. This allows the agent to make better decisions by knowing exactly which constraints are verified and which remain unverified.
Bare Assertion
This failure pattern occurs when an agent claims a constraint is satisfied without providing any supporting evidence from the search results. It means the agent asserts something as true without backing it up with verifiable data, leading to incorrect conclusions.

Terminology

Summary

Recent search agents often fail to reliably reason across all requirements in multi-constraint problems because they frequently engage in illusory completion, where they believe tasks are complete despite unresolved or violated constraints, leading to underverified answers.

The gist

Illusory completion frequently occurs wherein agents believe tasks are complete despite unresolved or violated constraints, leading to underverified answers.

How it works

The study investigates this capability under multi-constraint problems where valid answers must satisfy several constraints simultaneously. To diagnose illusory completion, the researchers introduced the EPISTEMIC LEDGER, an evaluation framework that tracks evidential support and agents’ beliefs for each constraint throughout multi-turn reasoning. This framework tracks two dimensions at each reasoning step: (i) whether the search results actually support each constraint, and (ii) whether the agent believes each constraint is satisfied. The ledger records evidential support E and agent belief B for each candidate-constraint pair, with E(k, Ci) ∈ SATISFIED, REFUTED, UNKNOWN and B(k, Ci) ∈ AFFIRM, DENY, UNADDRESS.

Illusion Diagnosis

The analysis of the EPISTEMIC LEDGER revealed four recurring failure patterns in search agents (Figure 2): (i) bare assertion, where an agent claims a constraint is satisfied without supporting evidence in the search results; (ii) overlooked refutation, where the agent ignores disconfirming evidence; (iii) stagnation, where the agent becomes stuck performing redundant searches that yield no new information on unverified constraints, leading to termination without resolution; and (iv) premature exit, where the agent terminates without ever addressing at least one constraint. These patterns suggest that illusory completion may stem from a lack of structured awareness of what has been verified and what remains unverified across candidates.

Mitigation Strategy

Motivated by these findings, the researchers examined whether explicit constraint-state tracking during execution mitigates these failures via LIVELEDGER, an inference-time tracker. This simple intervention consistently improves performance, substantially reducing underverified answers (by up to 26.5%) and improving overall accuracy (by up to 11.6%) on multi-constraint problems.

Experimental Setup

The experiments utilized 215 multi-constraint question-answering instances collected from five recent benchmarks, including BrowseComp, DeepSearchQA, FRAMES, LiveDRBench, and WebWalkerQA. The evaluation involved a diverse set of search agents spanning models trained via supervised or reinforcement learning and prompt-based methods. Metrics reported included accuracy (Acc), defined as the proportion of instances answered correctly, and underverified answer rate (UAR), defined as the proportion of instances in which the agent terminates with an underverified answer.

Key Findings

The results indicated that illusory completion prevails even among high-accuracy search agents, with UAR exceeding 50% for models like WebExplorer and TongyiDR. Illusory completion persists even within correct answers, as underverified correct answers still remain high, reaching 19.1% for TongyiDR. The EPISTEMIC LEDGER consistently increases accuracy and decreases UAR across all models, with the largest reduction of 27.5 points observed for ReAct-L-120B. LIVELEDGER exhibits the strongest synergy with the larger model, achieving substantial reductions across three failure patterns when using a 120B backbone.

Conclusion

LIVELEDGER consistently increases accuracy and decreases UAR across all models, demonstrating that explicit constraint checking improves both answer accuracy and reliability. The work concludes that explicit constraint-state information enables smoother and more efficient convergence to fully verified correct answers by steering agents away from invalid candidates. The paper suggests that while LIVELEDGER can mitigate illusory completion, there remains room for improvement in error-free constraint checking and in the digestion of ledger updates by search agents.

How it works

The EPISTEMIC LEDGER updates the candidate constraint state by jointly modeling evidential support and agent belief. The framework operates incrementally along the agent’s trajectory HT, updating evidence based on observations and beliefs based on subsequent reasoning traces. LIVELEDGER functions as an inference-time updater that exposes realtime updates on constraint satisfaction during execution by updating only the evidence support E(k, Cj) based on the observation ot. This augmentation allows the search agent to generate its next reasoning trace conditioned on the augmented context Ht,Lt+1∼π(· Ht,Lt+1).

Efficiency and Exploration

LIVELEDGER enables broader candidate exploration by quantifying behavior using the Extent of Candidate Exploration (ECE), which measures the average number of distinct candidates explored per reasoning turn. This behavior allows the agent to focus on candidates that genuinely have the potential to be the final answer, aligning with intended search behavior. Furthermore, LIVELEDGER makes search more efficient by reducing the number of turns taken in trajectories for agents like TongyiDR-L-20B. This suggests that explicit constraint-state information enables smoother and more efficient convergence to fully verified correct answers, by steering agents away from invalid candidates and toward verified ones. The paper plans to leverage the Epistemic Ledger to collect high-quality agent trajectories for post-training optimization.

Failure Patterns

The four recurring failure mechanisms identified are bare assertion, overlooked refutation, stagnation, and premature exit. LIVELEDGER consistently improves premature exits across all evaluated models, showing a reduction from 60% to 33% for TongyiDR. However, the effect of LIVELEDGER depends on the backbone model’s capacity to support explicit constraint checking and the search agent’s ability to accommodate such constraints. For instance, in ReAct-L-20B, most failures transition to None after integration, suggesting that truncated web snippets sometimes contain relevant keywords but omit critical contextual details, which can confuse the agent and lead to bare assertion errors.

Verification of Correct Answers

Manual review of verified but incorrect answers identified four major categories: correct answers missing from the ground truth, reasoning errors, reliance on unreliable or outdated information, and evaluation misclassification. The low rate of evaluation misclassification supports the robustness of the EPISTEMIC LEDGER framework, while the high proportion of verified-but-incorrect cases highlights the effectiveness of LiveLedger and underscores the importance of intermediate process evaluation alongside end-to-end evaluation. The study concludes that explicit constraint checking improves both answer accuracy and reliability.

Qualitative Examples

The paper provided qualitative examples illustrating incorrect agent execution (Figure 11) versus correct execution with LIVELEDGER (Figure 12), showing how the ledger tracks constraint satisfaction throughout the reasoning process. The example demonstrates that while an incorrect agent might incorrectly mark constraints as satisfied, the LIVELEDGER correctly identifies violations, leading to a fully verified correct answer.

Prompt Engineering

A specific prompt was developed for question decomposition that instructs the model to construct a Directed Acyclic Graph (DAG) of entities and dependency relations while extracting explicit constraints. This structured representation is crucial for identifying questions requiring the merging of three or more independent constraints to identify the answer.

Constraint Extraction

The paper introduced a prompt designed to extract explicit, externally verifiable constraints from a question, ensuring that any valid answer must satisfy these objectively verifiable conditions. This process decomposes the question into atomic, non-overlapping conditions that cannot be further decomposed into smaller conditions.

Ledger Updates

The EPISTEMIC LEDGER framework is implemented through two distinct prompts governing structured updates of the ledger: one for objective evidence and another for perception and status updates. The Objective Evidence Ledger Annotator's role is to update ‘null‘ values of ‘obj’ and ‘obj evidence‘ using ONLY the Search Results, ensuring that 'obj' reflects what has been objectively found in the Search Results. The Perception & Status Ledger Annotator updates candidate ‘status‘ and each constraint’s ‘per’ and ‘per evidence‘ using ONLY the agent’s thinking to capture its subjective beliefs.

Improvements for AI systems

  1. Bold header: Explicit Constraint Tracking via LIVELEDGER

This improvement involves integrating an inference-time updater, LIVELEDGER to expose real-time updates on constraint satisfaction during execution. This directly addresses the core failure of illusory completion, which occurs when agents incorrectly believe the query is fully resolved when some constraints remain unverified or violated.

  1. Bold header: Systematic Diagnosis of Failure Modes

By introducing the EPISTEMIC LEDGER, which tracks both evidential support and agents’ beliefs for each constraint, the system can diagnose four recurring failure patterns: (i) bare assertion, (ii) overlooked refutation, (iii) stagnation, and (iv) premature exit. This moves beyond simple final answer correctness to reveal epistemic failures invisible despite correct final answers.

  1. Bold header: Mitigation of Illusory Completion

The LIVELEDGER intervention is shown to be effective in reducing underverified answers by up to 26.5% and improving overall accuracy (by up to 11.6%) on multi-constraint problems. This demonstrates that explicit constraint-state tracking during execution mitigates these failures by forcing the agent toward exhaustive verification rather than premature commitment.

Sources

Related papers