ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "ReSI: Recursive Safety Improvement toward Resistant and Resilient AI".
Elias: The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated research,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So we're diving into "ReSI: Recursive Safety Improvement toward Resistant and Resilient AI," which lays out this recursive safety improvement approach.
Elias: The core idea here is that we need to adapt our safety alignment every time a model gets a new checkpoint because the safety established for one version doesn't necessarily carry over to its successors.
Priya: The paper claims that evolving red-teaming methods keep exposing new vulnerabilities in aligned language models, and this necessitates a continual adaptation of the safety alignment for each successive checkpoint.
Nadia: They term these goals resistance to known threats and resilience to unforeseen risks across all these checkpoints.
Elias: The ReSI framework organizes safety improvement as a recursive loop over target models, starting with an initial model M0, where in every round t they apply diverse red-teaming methods to find vulnerabilities in the current model Mt.
Priya: In each round, they develop a training recipe through automated research and then produce the next model Mt+one <ref:2610.12233#pg1>.
Nadia: The process involves three stages: Stage one collects successful attacks against Mt, Stage two screens combinations of attack sources and alignment methods through low-budget pilot trials, and Stage three conducts full-scale training.
Elias: The recipe itself is represented by a notation rho that specifies the contributions of different red-teaming sources, the alignment procedure, the training data mixture, and hyperparameters.
Priya: That data mixture has to balance learning from newly identified vulnerabilities with retaining existing safety behavior and general capabilities through proportions λcurrent plus λreplay plus λgeneral equaling one.
Nadia: The paper shows this method achieves in-distribution safety improvements on WildJailbreak and retention of existing defenses on HarmBench, outperforming A3 and matching or surpassing MAGIC in nearly all comparisons.
Elias: They also show this improvement extends to H-CoT and X-Teaming, challenging attack mechanisms that the model wasn't even trained on during the initial training.
Priya: The authors are showing that ReSI largely preserves reasoning, knowledge, instruction-following capabilities, and benign compliance even while improving safety.
Nadia: It’s about using experimental feedback to select and combine alignment methods while assessing safety improvements alongside instruction following and over-refusal for the next model update.
Elias: This whole paper is about building a validated model update as the target for the next round of improvement.
Priya: It supports automated experimentation as a practical approach to learning from newly discovered vulnerabilities in AI systems.
Conclusion: Nadia: We’ve walked through the details of ReSI, which is this recursive safety improvement framework that applies diverse red-teaming methods to identify vulnerabilities in the current target model.
Elias: It functions by developing training recipes through automated research and adopting a validated model update as the target for the next round of improvement.
Priya: Ultimately, this approach strengthens resistance to identified attacks while retaining existing defenses and demonstrates resilience to unseen risks.
Nadia: The authors have shown this framework can achieve resistance across different models and even resilience against unseen attacks on X-Teaming where four ReSI-trained models substantially outperform every frontier model comparator.
Elias: It’s a way to systematically tackle safety by constantly feeding the system new threats and refining its defenses in a bounded, operational sense.
Priya: This framework suggests that effective safety updates need training strategies suited to both the available attack data and the model's existing safety behavior for real-world deployment.
Nadia: ReSI is a recursive safety improvement framework that applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes through automated research, and adopts a validated model update as the target for the next round.
Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, Tat-Seng Chua, Xia Hu, Bowen Zhou, Chaochao Lu†♯],
Shanghai AI Laboratory · National University of Singapore · Shanghai Jiao Tong University · University of Science and Technology of China · Tsinghua University
cs.CR, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/MoonshotAI/Kimi-K3
License: http://creativecommons.org/licenses/by/4.0/
The gist: The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated
Key concepts
- Recursive Safety Improvement
- This is a loop where the system repeatedly applies safety enhancements. It starts with a base model, tests it for weaknesses using various attack methods, creates a better training strategy based on those findings, and then uses that strategy to create the next, safer version of the model. This cycle continues indefinitely.
- Red-Teaming Assessment
- This stage involves using many different adversarial testing techniques to try and break the current AI model. The goal is to find specific ways the model can be tricked into giving unsafe or harmful responses in a real-world setting. The results of these tests are recorded to identify exact vulnerabilities.
- Training Recipe Search Space
- A training recipe is a specific set of instructions for how to train the model next. ReSI searches through many combinations of attack sources and alignment methods to find the best recipe that strengthens safety without harming general abilities. It balances learning from new attacks with keeping existing safe behaviors.
- Pareto Gate
- This is a selection criterion used in Stage 3 to choose the final model update. A recipe passes the Pareto gate if it achieves significant safety improvements while simultaneously maintaining good instruction following and limiting excessive refusal compared to the original model, ensuring a balanced upgrade.
Terminology
Summary
The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated research, and promotes an update as the next target model.
The ReSI Framework
ReSI organizes safety improvement as a recursive loop over target models, starting from an initial model M0 (pg5). Within each round t, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current model Mt, develops a training recipe, and produces the next model Mt+1 (pg2). Each round follows three stages: Stage 1 collects successful attacks against Mt (pg4), Stage 2 screens combinations of attack sources and alignment methods through low-budget pilot trials (pg6), and Stage 3 conducts full-scale training and selects the update under the Pareto gate (pg7). The framework is recursive in a bounded, operational sense, where only the target model changes across rounds (pg4).
Recursive Improvement Process
The process involves searching an extensible space for a training recipe that strengthens resistance to exposed attacks while preserving existing safety behavior and general capabilities (pg6). A candidate recipe is represented as rho = (α, m, μ, h) which specifies the contributions of different red-teaming sources, the alignment procedure, the training data mixture, and hyperparameters (pg6). The training data mixture balances learning from newly identified vulnerabilities with retaining existing safety behavior and general capabilities through proportions λcurrent + λreplay + λgeneral = 1 (pg7).
Stage 1: Red-Teaming Assessment
In Stage 1, ReSI draws on a range of red-teaming methods to emulate adversarial interactions in deployment, probing for vulnerabilities that can lead to unsafe model responses (pg2). The framework aggregates records R t where the attack elicits an unsafe response (pg2). The attack dataset Dattack t captures vulnerabilities in the current model and provides safety training signals to guide the subsequent auto-research process (pg4).
Stage 2: Joint Recipe Search Space
Stage 2 involves screening combinations of attack sources and alignment methods through pilot trials (pg6). Initial proposals are constructed by pairing each attack source a ∈ A t with each alignment method m ∈ B, starting from an initial configuration where the mixture μ(0) = (1, 0, 0) (pg8). The resulting safety gain is measured as gt(ρ) = ASR(Mt; Vt) − ASR(Mρt; Vt) (pg11). Candidates are ranked by decreasing gt(ρ), and up to K recipes proceed to Stage 3 (pg6).
Stage 3: Full-Scale Refinement and Model Update
In Stage 3, ReSI conducts full-scale trials where the training data is expanded to all successful attacks obtained by applying method a to Mt over the full safety data pool S t (pg8). The Pareto gate requires safety improvement while preserving instruction following and limiting over-refusal relative to M0 (pg9). Eligible recipes are ranked by decreasing gt(ρ) for final selection, and the target model update follows Mt+1 = (Mρt, Et ≠ ∅, Mt, Et = ∅) (pg9).
Key Findings
ReSI improves in-distribution safety on WildJailbreak and retains or strengthens existing defenses on HarmBench, outperforming A3 and matching or surpassing MAGIC in nearly all available comparisons (pg12). These gains extend to H-CoT and X-Teaming, challenging attack mechanisms unseen during training that we use to assess resilience (pg12). On X-Teaming, all four aligned models outperform every evaluated frontier model, with ASRs of 16.98–43.40% versus 56.69–84.85% (pg12). ReSI largely preserves reasoning, knowledge, instruction-following capabilities, and benign compliance (pg12).
Contributions
The main contributions include the recursive safety improvement framework that recursively improves model safety by applying diverse red-teaming methods to identify vulnerabilities in the current target model (pg13). It achieves resistance to known attacks across four dense and mixture-of-experts models spanning different initial safety levels, matching or outperforming A3 and MAGIC on both WildJailbreak and HarmBench (pg13). Furthermore, it demonstrates resilience to unseen attacks on X-Teaming, where all four ReSI-trained models substantially outperform every frontier model comparator (pg13).
Discussion
These findings support automated experimentation as a practical approach to learning from newly discovered vulnerabilities (pg15). Effective safety updates require training strategies suited to both the available attack data and the model’s existing safety behavior (pg15). ReSI uses experimental feedback to select, combine, and refine alignment methods, assessing safety improvements alongside instruction following and over-refusal (pg15). The observed generalization to unseen attacks supports the value of this approach (pg15).
Conclusion
ReSI is a recursive safety improvement framework that applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes through automated research, and adopts a validated model update as the target for the next round (pg15). The framework strengthens resistance to identified attacks while retaining existing defenses and demonstrates resilience to unseen risks (pg15).
References
[1] Youbang Sun et al. R2AI: Towards Resistant and Resilient AI in an Evolving World. 2025. arXiv: 2509.06786 (pg1).
[3] Roman V. Yampolskiy. From Seed AI to Technological Singularity via Recursively Self-Improving Software. 2015. arXiv: 1502.06512 (pg4).
[7] Chris Lu et al. Towards end-to-end automation of AI research in Nature 651.8107 (2026) (pg3).
[9] Yueh-Han Chen, Jiaxin Wen, and Jan Hendrik Kirchner. Automated Researchers Can Mitigate Well-Characterized Alignment Failures. Alignment Science Blog (pg5).
[13] Xiaoyu Wen et al. JailbreakSkill: Scaling Automated Red-Teaming with Reusable and EverEvolving Skills (pg17).
[24] Youbang Sun et al. Position: Safe AI Should be Resistant and Resilient in an Evolving World. In: Forty-third International Conference on Machine Learning Position Paper Track (pg1).
[29] Xiaoyu Wen et al. MAGIC: A co-evolving attacker-defender adversarial game for robust LLM safety (pg17).
[30] Liwei Jiang et al. WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models. 2024. arXiv: 2406.18510 (pg18).
[37] Qiying Yu et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale (pg19).
[47] Ming Zhang et al. Safin-1: Safety from Within through Memory-Native State Evolution. 2026. arXiv: 2609.00092 (pg18).
[57] Yuntao Bai, Andy Jones, Kamal Ndousse, et al. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback (pg18).
[61] Melody Y. Guan et al. Deliberative Alignment: Reasoning Enables Safer Language Models (pg20).
[67] Anselm Paulus et al. Safety Alignment of LMs via Non-cooperative Games (pg20).
[73] Alexander Novikov et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery (pg21).
--- Page 1 ---
The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated research, and promotes an update as the next target model.
Stage 1: Red-Teaming Assessment
In Stage 1, ReSI draws on a range of red-teaming methods to emulate adversarial interactions in deployment, probing for vulnerabilities that can lead to unsafe model responses (pg2).
Improvements for AI systems
- Bold header: Recursive Safety Improvement Framework (ReSI)
The ReSI framework implements recursive self-improvement
by applying diverse red-teaming methods to identify vulnerabilities in the current target model, develops a training recipe through auto-research, and promotes a validated update as the next target model.
This allows AI systems to continually adapt their safety alignment based on newly exposed attack patterns.
- Bold header: Resistance to Known Threats
ReSI achieves resistance to known attacks
by selecting recipes that maximize safety gains while preserving instruction following and limiting over-refusal,
ensuring that updates do not regress on established safe behaviors relative to the initial model, as per the Pareto gate criteria.
- Bold header: Resilience to Unseen Attacks
The system gains resilience through generalization, as ReSI strengthens target-model safety against these outof-distribution attacks
by learning from new vulnerabilities and successfully challenging attack mechanisms unseen during training,
which is evidenced by performance on X-Teaming.
- Bold header: Adaptive Red-Teaming Integration
The framework incorporates extensible pools of red-teaming and alignment methods, allowing the loop to incorporate new attacks and learning approaches,
enabling the system to continuously expand its vulnerability identification capabilities.
- Bold header: Automated Recipe Search Space
ReSI searches a joint recipe space defined by attack data, alignment methods, training data mixtures, and hyperparameters—specifically selecting contributions from attack-data
(fractional selection of sources), alignment methods
(e.g., A3, SInternal), and training data mixtures
(balancing current safety examples with general-task examples).
- Bold header: Feedback-Guided Refinement Skills
The system refines its updates through specific skills, such as S1. Attack-source composition,
which expands source selection for limited gains, and S2. Harmful–benign pairing,
which constructs A3-style benign counterparts to harmful examples
to address retention failures in benign compliance.
- Bold header: Post-SFT Safety GRPO Continuation
The final refinement stage utilizes a safety GRPO process where the selected recipe's checkpoint is used, allowing for up to Rmax refinement trials
guided by conditions such as Safety gain ≥ 5 pp ASR reduction from Mt on the round’s safety metric.
Abstract
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R squared AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.
Sources
- \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World
- From Seed AI to Technological Singularity via Recursively Self-Improving Software
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Recursive self-improvement of AI research agents
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Jailbroken: How Does LLM Safety Training Fail?
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
- EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
- MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
- GPT-Red: Automated Red Teaming via Self-Play at Scale
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
- SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law
- Internalizing Safety Understanding in Large Reasoning Models via Verification
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs