Safety Must Survive Self-Improvement: Why Failures Persist and How Agents Recover

arXiv:2610.01073 · cs.SE, cs.CR · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Safety Must Survive Self-Improvement".

Nadia: Recursive self-improvement (RSI) allows agents to carry useful changes across generations, making it crucial to maintain safety by preventing unsafe behavior from persisting and enabling recovery when failures occur.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: Let’s talk about who put this paper together; it was written by Yunbei Zhang, Janet Wang, Saiyue Lyu, Yingqiang Ge, Kaiqu Liang, Zijian Jin, and Chandan K. Reddy. Their background is clearly deep in the fields of security research and agent development.

Elias: I noticed their affiliations span a few different universities across the US and Canada; that suggests a collaborative effort pulling expertise from several strong AI safety corners.

Priya: As someone focused on privacy, I wonder if having researchers from different institutional backgrounds helps ensure the testbed they built for this study is as robust as possible against unforeseen edge cases in authorization tasks.

Nadia: That’s true; when you're dealing with stateful authorization—things like session authorization or tool approval—you need diverse perspectives to spot where a simple fix might introduce a hidden vulnerability later on.

Elias: I think the implication here is that for any system relying on recursive self-improvement, the safety mechanism can’t just be an initial filter; it has to account for how those changes ripple through the entire history of the agent's decisions.

Priya: So, when we look at this paper, we’re not just looking at one specific vulnerability; we are looking at a pattern of failure persistence in complex, evolving AI systems.

Nadia: Precisely; it moves us past thinking about isolated bugs and into the long-term stability of autonomous agents as they iterate on their own code.

Elias: And that leads us directly into what the paper actually claims is happening during this self-improvement process, which we’ll get to in a moment.

The paper's summary: Nadia: So, the core of "Safety Must Survive Self-Improvement: Why Failures Persist and How Agents Recover" is that recursive self-improvement allows agents to carry useful changes across generations, but maintaining safety requires preventing unsafe behavior from sticking around and having a way to recover when things inevitably go wrong.

Elias: It’s summarizing how the agent inherits not just the capabilities of its predecessors but also the underlying assumptions those predecessors made about safety, which is a big thing for any cryptographic system we look at.

Priya: The summary explains that they used a controlled testbed with four stateful authorization families—session authorization, filesystem containment, tool approval, and structured user consent—interleaved with event streams that mix requests with scope changes.

Nadia: That setup is crucial because it lets them observe exactly what happens when an agent tries to optimize itself while simultaneously facing real-world requests that might require a different set of permissions.

Elias: They found that when they look at the historical scores, in twenty-two out of forty-eight framework histories, the same unsafe programs stayed active even though there was a correct alternative available in every affected archive.

Priya: That specific finding is really telling because it points directly to the issue of historical eligibility preserving dangerous code despite better options existing within the past versions.

Nadia: And that persistence happens because of two main mechanisms they identified: historical eligibility and keep-after-rejection, which allows a failed program to keep running even when all proposed fixes fail validation.

Elias: So, the summary boils down to showing that evaluation alone doesn't guarantee safe execution continuity once a failure is observable during the agent's evolution.

The paper's improvements: Nadia: The paper suggests several ways we can improve how these agents handle safety during their optimization loop, and it proposes looking at what executes and which code supplies the next edit as the primary levers for recovery.

Elias: They compare four different strategies—checking the current program, following a policy, running passing code, or refreshing archive scores—to see which one actually leads to a correct outcome faster.

Priya: The study highlights that while fully validating and rolling back can lead to fully correct programs in the core trajectory study, it also saves over forty-three percent in deployment costs because we don't always need the absolute most perfect version immediately.

Nadia: That’s a practical point; we can get a program that is safe enough for deployment while still gaining significant efficiency compared to waiting for perfect validation.

Elias: Furthermore, they found that editing the initial implementation helps recovery from shared failures relative to editing the failed version, though this advantage changes depending on which specific LLM editor you are using.

Priya: This suggests that where we start our optimization process matters significantly for how resilient the agent is when it encounters failures across multiple related tasks.

Nadia: So, the improvement isn't just about finding a fix; it’s about designing a dynamic decision process—choosing which program runs and which code gets edited next—to keep safety intact across generations.

Conclusion: Nadia: To wrap things up on "Safety Must Survive Self-Improvement: Why Failures Persist and How Agents Recover," the authors conclude that execution and editing decisions jointly determine whether useful improvements preserve safety across generations, emphasizing the need to check what will run under current conditions and choose which implementation to edit next.

Elias: It seems the central message is that safety must be maintained throughout every stage of improvement, even if a revision isn't accepted by the validator.

Priya: From my view, this paper really underscores that when we look at these complex evolving systems, we can't rely on a single check; we need this joint decision-making between execution and editing choices to ensure safety is preserved.

Nadia: It’s a strong argument for building more nuanced control mechanisms into the agent architecture itself rather than relying solely on post-hoc validation of its output.

Elias: I think this work provides concrete evidence of where the theoretical concerns about iterative system evolution actually manifest in real, stateful authorization tasks.

Priya: It gives us a framework for how to measure not just correctness, but also the cost and reliability associated with maintaining safety during that constant optimization process.

Nadia: We’ve covered a lot of ground on this paper today, showing exactly why we need to be careful about letting agents optimize themselves too freely.

Yunbei Zhang, Janet Wang, Saiyue Lyu, Yingqiang Ge, Kaiqu Liang, Zijian Jin, Chandan K. Reddy

Tulane University · University of British Columbia University of British Columbia Department of Computer Science and Engineering Research Institute Rutgers University Princeton University New York University Virginia Tech

cs.SE, cs.CR

Submitted: 2026-10-01

Updated: 2026-10-01

Comments: 44 pages, 13 figures. Project page: https://RSI-Safety.github.io/

Code: https://github.com/apache/casbin-pycasbin

Project page: https://rsi-safety.github.io/ABSTRACT

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Recursive self-improvement (RSI) allows agents to carry useful changes across generations, making it crucial to maintain safety by preventing unsafe behavior from persisting and enabling recovery

Key concepts

Recursive Self-Improvement (RSI)
RSI is when an AI agent repeatedly optimizes its own code across generations. The study uses this to see if beneficial changes can be made safely, but it highlights the danger of unsafe behaviors sticking around.
Historical Eligibility
'Historical scores' are metrics that track past performance. The study found these scores often favor an unsafe program over a correct fix, meaning the system remembers and keeps old, risky code even when better options exist.
Keep-after-rejection
'Keep-after-rejection' is a mechanism where if the current program and all proposed changes fail validation, the system keeps the original failed program active. This allows unsafe execution to continue even when all alternatives are rejected.
Editing Source
This refers to whether an agent starts its optimization from scratch (founder editing) or edits an existing version. The study showed that starting with a fresh implementation improves recovery from shared failures compared to trying to fix the failed version directly.

Terminology

Summary

Recursive self-improvement (RSI) allows agents to carry useful changes across generations, making it crucial to maintain safety by preventing unsafe behavior from persisting and enabling recovery when failures occur. This study investigates these challenges through a controlled testbed of stateful authorization tasks, demonstrating that safety depends on what runs and what supplies the next edit.

How it works

The research constructs a controlled testbed featuring four stateful authorization families—session authorization, filesystem containment, tool approval, and structured user consent—interleaved with seeded event streams that interleave requests with revocation or scope changes. Fixed LLM editors repeatedly optimize executable agent components while selected programs and their histories evolve. Independent traces distinguish unauthorized effects from the completion of permitted work. The core mechanism involves an improvement loop where at generation g, the incumbent is wg−1, and the editor produces two children u(1)g and u(2)g based on a parent sg.

Failure Persistence Mechanisms

The study identifies two primary mechanisms through which failures persist despite correct alternatives:

  1. Historical eligibility: Historical scores preserve the same unsafe programs in 22 of 48 framework histories despite a correct alternative in every affected archive. This occurs when historical scores favor an unsafe program over a correct repair.

  2. Keep-after-rejection: when the current program and all proposals fail validation, keep-after-rejection leaves the failed current program active. This mechanism allows an unsafe execution to continue even when all candidates fail validation.

Recovery and Control Mechanisms

The paper examines distinct outcomes governing recovery:

  1. Execution and editing source govern recovery: Which program executes? Which code supplies the next edit? The study compares four strategies: Check current, policy, Run passing code, Refresh archive scores.

  2. Validation and validated rollback end with fully correct programs in the core trajectory study while retaining over 43% deployment savings.

  3. Editing source affects recovery from shared failures: Editing the initial implementation improves recovery from shared failures relative to editing the failed version, although the advantage varies across editors.

Key Findings on Control Strategies

The controlled comparisons separate repair availability from deployment decisions:

. Full initial validation does not prevent later failure under new authorization dependencies.

. Refreshing scores restores correctness on the original suite, while independently composed tests expose remaining coverage gaps.

. Under an unchanged contract, matched comparisons confirm that keeping the incumbent after rejection can preserve a detected failure.

Impact of Editing Source and Utility

The study isolates the effect of editing source on recovery from shared failures: Founder editing improves recovery from shared failures relative to editing the failed version, with the advantage varying across editors. Furthermore, it investigates whether preserving safety still permits useful optimization: Safety alone does not imply full correctness. Validated rollback and full validation finish with fully correct endpoints while retaining mean service-cost savings above 43%.

Conclusion

The findings suggest that execution and editing decisions jointly determine whether useful improvements preserve safety across generations, requiring checks on what will run under current conditions and choosing which implementation to edit next. Safety must be preserved throughout improvement, even when no revision is accepted.

Limitations

The study assumes correct validators within the tested conditions, and larger deployments or jointly evolving editors remain directions for further study. The main controls assume a correct founder, an externally specified contract, and a correct validator. The all-history-invalid control tests restoration without a valid founder. Independent compositions expose coverage gaps. Service costs measure weighted operations rather than model inference or latency.

References

(A comprehensive list of references is provided in Appendix I.)


(Self-Correction/Refinement Note: The request asks for the summary to be structured exactly like the provided template, which includes a specific The gist sentence and then 3 to 5 sections starting with bold headers. I will ensure the final output strictly adheres to this structure while maintaining the required length and tone.)

The gist

Safety depends on what runs and what supplies the next edit.

How it works

Recursive self-improvement (RSI) allows agents to carry useful changes across generations, making it crucial to maintain safety by preventing unsafe behavior from persisting and enabling recovery when failures occur. This study investigates these challenges through a controlled testbed of stateful authorization tasks, demonstrating that safety depends on what runs and what supplies the next edit. The core mechanism involves an improvement loop where at generation g, the incumbent is wg−1, and the editor produces two children u(1)g and u(2)g based on a parent sg.

Failure Persistence Mechanisms

The research identifies two primary mechanisms through which failures persist despite correct alternatives:

Improvements for AI systems

Here are the specific improvements that can be made to AI systems, based on the findings of this research:

  1. A system designed for recursive self-improvement (RSI) must incorporate a multi-layered safety mechanism that goes beyond simple validation checks. This mechanism must simultaneously monitor:

  2. The current execution eligibility against the most recent contract,

  3. Historical eligibility via stored scores/archives, and

  4. The implementation source used to generate the next revision (editing source).

This improved AI system can do the following:

  1. Maintain safety across generations of self-improvement by actively choosing between three distinct recovery strategies: current deployment validation, refreshing historical eligibility scores, or using a validated fallback mechanism (e.g., restoring a founder).

  2. Prevent unsafe behavior from persisting by continuously checking what will run under current conditions before selecting which implementation to edit next.

  3. Improve recovery from shared failures by intelligently choosing whether to edit the initial correct implementation versus the failed version, depending on the specific LLM editor and feedback context.

  4. Ensure that useful optimization gains (efficiency) are preserved even when safety is restored, by measuring service cost savings separately from full correctness metrics (authorized effects + completion of permitted work).

  5. Provide a system that can distinguish between a detected failure, the availability of a repair in the archive, and the decision to end unsafe execution—allowing controllers to choose whether to pause execution during repair or select an existing valid alternative.

Abstract

Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization tasks, where fixed LLM editors optimize executable agent components and independent traces record their effects. Paired interventions separate which revisions pass validation, which program continues running, and which program the editor revises next. After a new authorization dependency invalidates previously tested optimizations, historical scores preserve the same unsafe programs in 22 of 48 framework histories despite a correct alternative in every affected archive. Refreshing scores restores correctness on the original suite, with residual failures on independently composed tests. Failures also persist under an unchanged contract when all proposals are rejected and the failed incumbent remains active. Starting from shared failures, editing the initial implementation instead of the failed one improves recovery, although the advantage varies across editors. Full validation and validated rollback end with fully correct programs in the core trajectory study while retaining over 43% deployment savings. Preserving agent safety requires checking what will run under current conditions and choosing which implementation to edit next.

Sources

Related papers