Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
summary
The gist
This paper examines the gap between finding a correct code patch and successfully retaining, verifying, and submitting it within agentic loops.
In short
The paper "Looping Is Not Reliability" challenges the assumption that more AI agent attempts lead to better code repair. The hosts discuss how correctness can drop during iterations and propose "evidence-bound typed loop contracts" using specific revision actions and code hashes to ensure stability and prevent errors.
Key concepts
- Stale Evidence
- Stale evidence occurs when an AI agent makes decisions based on outdated information or traces that do not match the current state of the code. This mismatch can lead to a significant increase in errors, as the agent is essentially working with instructions meant for a different version of the software.
- Evidence-Bound Typed Loop Contract
- An evidence-bound typed loop contract is a proposed framework designed to force AI agents to follow strict, predictable rules during revisions. By tying feedback to specific code hashes and using structured commands, this method prevents agents from making mistakes based on outdated information and ensures every action is verifiable.
- Typed Revision Actions
- Typed revision actions are specific, structured commands—such as "Keep," "Patch," or "Escalate"—that an agent uses instead of writing free-form text. This approach prevents the agent from producing unparsable code or becoming stuck, providing a disciplined and auditable framework for autonomous software repair processes.
Terminology used across episodes
This episode discusses
- Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair · Paper Radio
- Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows · Paper Radio
- Evaluating Large Language Models Trained on Code
- FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback
- EviACT: An Evidence-to-Action Framework for Agentic Program Repair
- AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
- Qwen2.5 Technical Report
- AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents · Paper Radio
The paper
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair · Read on arXiv
Alibaba Cloud
Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95% CI [8.9,37.0], exact Holm p=0.0337). A prospective 540-rollout policy eliminates observed correct-start harm but reduces wrong-start repair and fails its joint criterion. Repository experiments over 24 bugs and four coder stacks expose floor effects and component heterogeneity without Holm-significant effects. We therefore separate admission, preservation, grounded certification, competence, and liveness. We derive an evidence-bound typed loop contract and instantiate its mechanically enforceable subset in a reference implementation that binds verifier evidence to exact code states, preserves verified checkpoints, and emits auditable admission receipts. The implementation is an executable specification and conformance artifact, not evidence of improved repair competence or calibrated verifier dependence.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair".
Jane: The paper was written by Xueping Gao, Jianwei Yang and Qiang Yang from Alibaba Cloud.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, I’ve been reading through this paper "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair," and it feels like a massive reality check for the whole industry.
Jane: It really does, Tom. The authors—Xueping Gao, Jianwei Yang, and Qiang Yang from Alibaba Cloud—are essentially pulling back the curtain on how these coding agents actually function during their iterative loops.
Tom: They’re challenging that basic idea we all have that if an AI just tries a task over and over again, it will eventually get it right.
Jane: Exactly, because they've discovered that more attempts can sometimes actually lead to more mistakes rather than more success.
Lu: I think the implications are quite profound for how we view machine intelligence. We aren't just looking at a lack of smarts, but a lack of formal structure in how these agents interact with their environment.
Meng: That sounds like a nightmare for anyone trying to build dependable software. If we can't trust the loop, we can't trust the agent to work on our real-world codebases without constant supervision.
Jane: Meng is right, and that’s why the paper is so important; it moves us away from just "hoping" for success toward a more disciplined approach.
Lu: We could see a future where agents operate with a level of mathematical certainty that we've never seen before, almost like they have their own internal laws.
Lalam: It represents a shift in our cultural expectation of technology. We are moving from treating AI as this unpredictable magic box to seeing it as a professional tool that must follow strict protocols to be useful.
Tom: That brings us directly to the data they gathered, which is honestly quite startling when you see the numbers.
Summary: Tom: We’ve been talking about the risks, but let's look at what actually happens in these loops according to "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair."
Jane: The researchers found that while an agent might find a correct solution at some point, it often loses that correctness in the very next step.
Tom: I saw those figures, Jane; they noted that current correctness can actually drop from eighty-two percent after one revision down to sixty-seven point three percent after just two revisions.
Jane: It’s a bit of a rollercoaster, isn't it? You think you've reached the finish line, but the agent just trips over its own feet on the next attempt.
Meng: That is a massive problem for stability. If I'm an engineer and my agent breaks a working fix while trying to "improve" it, that's time and money wasted every single time.
Tom: And it gets even worse when they introduced what they call "stale evidence."
Jane: That was one of the most interesting parts of the study. The 14B model saw a huge twenty-two point two-point increase in harm when it was given outdated information instead of current traces.
Lu: It's like trying to fix a modern electric car using a manual for an old steam engine! The information might be technically correct for *some* machine, but it’s completely wrong for the one in front of you.
Meng: If the agent is basing its decisions on data that doesn't match the current state of the code, we aren't even doing engineering anymore.
Lalam: This shows that intelligence alone isn't enough for an agent to be reliable. It must be perfectly synchronized with the reality of the task it is performing.
Tom: So, how do we stop this cycle of making things worse?
Improvements: Tom: The authors aren't just pointing out flaws; they provide a specific blueprint in "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair" to fix these loops.
Jane: They propose something called an "evidence-bound typed loop contract." It's a way to force the agent to follow very specific rules every time it makes a change.
Tom: I was particularly interested in those "typed revision actions" they mentioned.
Jane: Instead of letting the agent just write whatever text it wants, they want it to use specific commands like "Keep," "Patch," or "Escalate."
Tom: That seems like it would prevent a lot of the confusion when an agent is stuck or produces unparsable code.
Meng: I really like the idea of tying every piece of feedback to a specific code hash. If you use hashes to bind evidence to the state, you basically kill that "stale evidence" problem they found earlier.
Lu: It’s as if they are giving the agent a digital notary! Every action is recorded and verified against the exact version of the code being worked on.
Meng: And they also mentioned keeping "last-known-good checkpoints," which is a standard practice we use in software development.
Jane: Exactly, so if a new revision fails, the system can just roll back to that last verified state instead of letting the agent wander off into a broken mess.
Lalam: This creates a sense of accountability. We are building systems that don't just act, but act within an auditable and predictable framework.
Tom: It's a complete redesign of how we manage these autonomous processes.
Conclusion: Tom: We have reached the end of our discussion on "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair."
Jane: This paper really changes how we need to think about agent evaluation. We can't just look at whether an agent eventually succeeds; we have to look at how stable it is throughout the whole process.
Lu: I'm thinking about the massive potential here for creating truly autonomous software factories that operate under strict, verifiable quality controls!
Meng: And from my side, it gives me a real technical path toward actually trusting these agents in my production environments.
Lalam: It's a step toward a culture where AI is seen as a disciplined and reliable partner in our most complex engineering tasks.
Tom: Thanks for joining us today, everyone! We'll see you next time with another fascinating paper.
Jane: Goodbye, everyone!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization