Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair".
Jane: The paper was written by Xueping Gao, Jianwei Yang and Qiang Yang from Alibaba Cloud.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, I’ve been reading through this paper "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair," and it feels like a massive reality check for the whole industry.
Jane: It really does, Tom. The authors—Xueping Gao, Jianwei Yang, and Qiang Yang from Alibaba Cloud—are essentially pulling back the curtain on how these coding agents actually function during their iterative loops.
Tom: They’re challenging that basic idea we all have that if an AI just tries a task over and over again, it will eventually get it right.
Jane: Exactly, because they've discovered that more attempts can sometimes actually lead to more mistakes rather than more success.
Lu: I think the implications are quite profound for how we view machine intelligence. We aren't just looking at a lack of smarts, but a lack of formal structure in how these agents interact with their environment.
Meng: That sounds like a nightmare for anyone trying to build dependable software. If we can't trust the loop, we can't trust the agent to work on our real-world codebases without constant supervision.
Jane: Meng is right, and that’s why the paper is so important; it moves us away from just "hoping" for success toward a more disciplined approach.
Lu: We could see a future where agents operate with a level of mathematical certainty that we've never seen before, almost like they have their own internal laws.
Lalam: It represents a shift in our cultural expectation of technology. We are moving from treating AI as this unpredictable magic box to seeing it as a professional tool that must follow strict protocols to be useful.
Tom: That brings us directly to the data they gathered, which is honestly quite startling when you see the numbers.
Summary: Tom: We’ve been talking about the risks, but let's look at what actually happens in these loops according to "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair."
Jane: The researchers found that while an agent might find a correct solution at some point, it often loses that correctness in the very next step.
Tom: I saw those figures, Jane; they noted that current correctness can actually drop from eighty-two percent after one revision down to sixty-seven point three percent after just two revisions.
Jane: It’s a bit of a rollercoaster, isn't it? You think you've reached the finish line, but the agent just trips over its own feet on the next attempt.
Meng: That is a massive problem for stability. If I'm an engineer and my agent breaks a working fix while trying to "improve" it, that's time and money wasted every single time.
Tom: And it gets even worse when they introduced what they call "stale evidence."
Jane: That was one of the most interesting parts of the study. The 14B model saw a huge twenty-two point two-point increase in harm when it was given outdated information instead of current traces.
Lu: It's like trying to fix a modern electric car using a manual for an old steam engine! The information might be technically correct for *some* machine, but it’s completely wrong for the one in front of you.
Meng: If the agent is basing its decisions on data that doesn't match the current state of the code, we aren't even doing engineering anymore.
Lalam: This shows that intelligence alone isn't enough for an agent to be reliable. It must be perfectly synchronized with the reality of the task it is performing.
Tom: So, how do we stop this cycle of making things worse?
Improvements: Tom: The authors aren't just pointing out flaws; they provide a specific blueprint in "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair" to fix these loops.
Jane: They propose something called an "evidence-bound typed loop contract." It's a way to force the agent to follow very specific rules every time it makes a change.
Tom: I was particularly interested in those "typed revision actions" they mentioned.
Jane: Instead of letting the agent just write whatever text it wants, they want it to use specific commands like "Keep," "Patch," or "Escalate."
Tom: That seems like it would prevent a lot of the confusion when an agent is stuck or produces unparsable code.
Meng: I really like the idea of tying every piece of feedback to a specific code hash. If you use hashes to bind evidence to the state, you basically kill that "stale evidence" problem they found earlier.
Lu: It’s as if they are giving the agent a digital notary! Every action is recorded and verified against the exact version of the code being worked on.
Meng: And they also mentioned keeping "last-known-good checkpoints," which is a standard practice we use in software development.
Jane: Exactly, so if a new revision fails, the system can just roll back to that last verified state instead of letting the agent wander off into a broken mess.
Lalam: This creates a sense of accountability. We are building systems that don't just act, but act within an auditable and predictable framework.
Tom: It's a complete redesign of how we manage these autonomous processes.
Conclusion: Tom: We have reached the end of our discussion on "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair."
Jane: This paper really changes how we need to think about agent evaluation. We can't just look at whether an agent eventually succeeds; we have to look at how stable it is throughout the whole process.
Lu: I'm thinking about the massive potential here for creating truly autonomous software factories that operate under strict, verifiable quality controls!
Meng: And from my side, it gives me a real technical path toward actually trusting these agents in my production environments.
Lalam: It's a step toward a culture where AI is seen as a disciplined and reliable partner in our most complex engineering tasks.
Tom: Thanks for joining us today, everyone! We'll see you next time with another fascinating paper.
Jane: Goodbye, everyone!
Alibaba Cloud
cs.CL, cs.AI
Submitted: 2026-07-27
Updated: 2026-10-04
Comments: 11 pages, 5 figures, 8 tables. Accepted at AgenticDev 2026, co-located with ASE 2026. Camera-ready version; incorporates presentation revisions and final publication metadata. Scientific conclusions are unchanged
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: This paper examines the gap between finding a correct code patch and successfully retaining, verifying, and submitting it within agentic loops.
Key concepts
- Stale Evidence
- Stale evidence occurs when an AI agent makes decisions based on outdated information or traces that do not match the current state of the code. This mismatch can lead to a significant increase in errors, as the agent is essentially working with instructions meant for a different version of the software.
- Evidence-Bound Typed Loop Contract
- An evidence-bound typed loop contract is a proposed framework designed to force AI agents to follow strict, predictable rules during revisions. By tying feedback to specific code hashes and using structured commands, this method prevents agents from making mistakes based on outdated information and ensures every action is verifiable.
- Typed Revision Actions
- Typed revision actions are specific, structured commands—such as "Keep," "Patch," or "Escalate"—that an agent uses instead of writing free-form text. This approach prevents the agent from producing unparsable code or becoming stuck, providing a disciplined and auditable framework for autonomous software repair processes.
Terminology
Summary
This paper examines the gap between finding a correct code patch and successfully retaining, verifying, and submitting it within agentic loops. It argues that current generate–test–revise
cycles lack completion reliability,
as repetition alone provides no guarantee that a correct state will not be destroyed by subsequent revisions or stale feedback.
The non-absorbing nature of correctness
The study reveals that correctness is not an absorbing state under repeated revision.
While the cumulative ever-correct
rate may increase, the current correctness
often declines during forced iterations. This phenomenon is driven by several factors:
-
Stale evidence, such as a failure trace from a previous version, can cause a coder to regress from a correct state.
-
The effectiveness of evidence depends on the
joint effective action
of the payload and the coder's responsiveness. -
Stale traces harm
correct starts by introducingstate-misalignment,
where valid facts are applied to the wrong code version, effectively resurrecting fixed bugs.
Verifier risk and orchestration costs
The researchers demonstrate that verifier quality is not independence,
meaning model-family diversity does not automatically ensure low risk or independent errors. They identify several critical limitations in current orchestration:
-
Conditional false-accept dependence
acts as an additive reliability penalty for ensembles, even if marginal verifiers are individually strong. -
Selective risk at a specified coverage
can remain high even when using multiple models, and changing the acceptance operating point changes which errors are accepted. -
Bundled admission guards often suffer from a
safety–liveness trade-off,
where reducing unsafe completions also reducesrevision plasticity and completion coverage,
failing to improve overall liveness in repository-scale settings.
The evidence-bound typed loop contract
To mitigate these risks, the authors propose an evidence-bound typed loop contract
that transforms permissive text loops into a structured state machine. This contract enforces several testable obligations:
-
State-bound evidence: Every observation must be carried in an envelope where the code hash matches the current state to prevent
state-misalignment.
-
Typed revision actions: Agents are restricted to a formal set of actions—
Keep,Patch(code), orEscalate(reason)—to prevent free-form text from conflatingno edit is needed
with an unparsable patch. -
Last-known-good checkpoint: The system must maintain a
verified checkpoint
that is protected from being overwritten by invalid actions or malformed patches. -
Risk-aware stopping: Completion decisions must be based on
fresh completion certification
at a declared risk–coverage operating point.
StateSeal implementation and conformance
The authors provide StateSeal,
a reference implementation that functions as an admission layer around existing coding agents and verifiers.
Rather than attempting to improve the agent's underlying repair competence, StateSeal serves as an executable specification and conformance artifact.
It mechanistically enforces the contract by:
-
Intercepting state-hash mismatches to prevent erroneous commits.
-
Ensuring that parser or action-schema failures cannot erase a previously verified checkpoint.
-
Providing
auditable admission receipts
that link verifier evidence to exact code states.
The implementation is designed so that a state-hash mismatch cannot trigger a commit
and parser or action-schema failure cannot erase a checkpoint,
ensuring these safety properties hold even when model behavior is unpredictable.
Improvements for AI systems
1. State-Bound Evidence Orchestration
-
Improvement: Implement a mandatory cryptographic binding between every piece of feedback (execution traces, PASS/FAIL status, error logs) and the specific SHA-256 hash of the code state (h code) and test suite (h suite) that produced it.
-
Capability: The system will automatically reject or trigger a
refresh
cycle if an agent attempts to apply feedback that does not match the current code's identity, effectively eliminatingstale trace regression
where agents attempt to fix bugs that have already been resolved or altered.
2. Typed Revision Action Schema
-
Improvement: Replace free-form text generation for agentic decisions with a strictly parsed, executable action schema:
Keep(no change),Patch(code)(new code proposal), orEscalate(reason)(abstain/request human intervention). -
Capability: The system will prevent
silent failure
loops where an agent produces unparsable text, empty patches, or non-functional status updates (e.g.,STATUS: DONE
) that would otherwise cause the orchestrator to erroneously overwrite a correct code state with an invalid one.
3. Last-Known-Good (LKG) Checkpoint Ledger
-
Improvement: Integrate an immutable ledger that records the code and evidence hashes of the most recent state to pass all configured verification gates (visible, challenge, and hidden tests).
-
Capability: The system will enable automatic
rollback
functionality. If a new revision fails a gate or introduces a regression, the orchestrator will mechanically revert to the LKG state rather than allowing the agent to continue iterating from a broken or incorrect baseline.
4. Dependence-Aware Verification Gate
-
Improvement: Transition from simple model-consensus stopping rules to a policy based on conditional false-accept dependence (phi) and risk at target coverage. This involves selecting verifiers not just for their individual accuracy, but for their lack of correlated error patterns.
-
Capability: The system will prevent
correlated error acceptance,
where multiple LLM verifiers confirm the same incorrect patch. It ensures the agent only terminates when the joint probability of error is mathematically below a pre-defined risk threshold, rather than relying on model diversity as a proxy for reliability.
5. Freshness-Certified Completion Protocol
-
Improvement: Implement an admission layer that requires
fresh certification
—re-executing all verification suites on the final proposed patch in an isolated environment before it is committed to the repository. -
Capability: The system will prevent
completion-gate dependence,
where a system accepts a patch based on stale or partially successful tests, ensuring that every submitted solution has been independently verified against the exact final state of the code.
Abstract
Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95% CI [8.9,37.0], exact Holm p=0.0337). A prospective 540-rollout policy eliminates observed correct-start harm but reduces wrong-start repair and fails its joint criterion. Repository experiments over 24 bugs and four coder stacks expose floor effects and component heterogeneity without Holm-significant effects. We therefore separate admission, preservation, grounded certification, competence, and liveness. We derive an evidence-bound typed loop contract and instantiate its mechanically enforceable subset in a reference implementation that binds verifier evidence to exact code states, preserves verified checkpoints, and emits auditable admission receipts. The implementation is an executable specification and conformance artifact, not evidence of improved repair competence or calibrated verifier dependence.
Sources
- Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
- Evaluating Large Language Models Trained on Code
- FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback
- EviACT: An Evidence-to-Action Framework for Agentic Program Repair
- AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
- Qwen2.5 Technical Report
- AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering