ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
summary
The gist
The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated
In short
The ReSI framework is a recursive safety improvement system that iteratively enhances model safety. It uses diverse red-teaming methods to find vulnerabilities, develops automated training recipes, and produces an updated target model in each cycle. This process aims to build resilient AI by continuously testing and improving defenses against attacks.
Key concepts
- Recursive Safety Improvement
- This is a loop where the system repeatedly applies safety enhancements. It starts with a base model, tests it for weaknesses using various attack methods, creates a better training strategy based on those findings, and then uses that strategy to create the next, safer version of the model. This cycle continues indefinitely.
- Red-Teaming Assessment
- This stage involves using many different adversarial testing techniques to try and break the current AI model. The goal is to find specific ways the model can be tricked into giving unsafe or harmful responses in a real-world setting. The results of these tests are recorded to identify exact vulnerabilities.
- Training Recipe Search Space
- A training recipe is a specific set of instructions for how to train the model next. ReSI searches through many combinations of attack sources and alignment methods to find the best recipe that strengthens safety without harming general abilities. It balances learning from new attacks with keeping existing safe behaviors.
- Pareto Gate
- This is a selection criterion used in Stage 3 to choose the final model update. A recipe passes the Pareto gate if it achieves significant safety improvements while simultaneously maintaining good instruction following and limiting excessive refusal compared to the original model, ensuring a balanced upgrade.
Terminology used across episodes
This episode discusses
- ReSI: Recursive Safety Improvement toward Resistant and Resilient AI · Paper Radio
- R squared AI: Towards Resistant and Resilient AI in an Evolving World
- From Seed AI to Technological Singularity via Recursively Self-Improving Software
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Recursive self-improvement of AI research agents · Paper Radio
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Jailbroken: How Does LLM Safety Training Fail?
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · Paper Radio
- EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking · Paper Radio
- MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
- GPT-Red: Automated Red Teaming via Self-Play at Scale · Paper Radio
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- Towards AI- 45 Law: A Roadmap to Trustworthy AGI
- SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law
- Internalizing Safety Understanding in Large Reasoning Models via Verification
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
The paper
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI · Read on arXiv
Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, Tat-Seng Chua, Xia Hu, Bowen Zhou, Chaochao Lu†♯],
Shanghai AI Laboratory · National University of Singapore · Shanghai Jiao Tong University · University of Science and Technology of China · Tsinghua University
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R squared AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "ReSI: Recursive Safety Improvement toward Resistant and Resilient AI".
Elias: The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated research,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So we're diving into "ReSI: Recursive Safety Improvement toward Resistant and Resilient AI," which lays out this recursive safety improvement approach.
Elias: The core idea here is that we need to adapt our safety alignment every time a model gets a new checkpoint because the safety established for one version doesn't necessarily carry over to its successors.
Priya: The paper claims that evolving red-teaming methods keep exposing new vulnerabilities in aligned language models, and this necessitates a continual adaptation of the safety alignment for each successive checkpoint.
Nadia: They term these goals resistance to known threats and resilience to unforeseen risks across all these checkpoints.
Elias: The ReSI framework organizes safety improvement as a recursive loop over target models, starting with an initial model M0, where in every round t they apply diverse red-teaming methods to find vulnerabilities in the current model Mt.
Priya: In each round, they develop a training recipe through automated research and then produce the next model Mt+one <ref:2610.12233#pg1>.
Nadia: The process involves three stages: Stage one collects successful attacks against Mt, Stage two screens combinations of attack sources and alignment methods through low-budget pilot trials, and Stage three conducts full-scale training.
Elias: The recipe itself is represented by a notation rho that specifies the contributions of different red-teaming sources, the alignment procedure, the training data mixture, and hyperparameters.
Priya: That data mixture has to balance learning from newly identified vulnerabilities with retaining existing safety behavior and general capabilities through proportions λcurrent plus λreplay plus λgeneral equaling one.
Nadia: The paper shows this method achieves in-distribution safety improvements on WildJailbreak and retention of existing defenses on HarmBench, outperforming A3 and matching or surpassing MAGIC in nearly all comparisons.
Elias: They also show this improvement extends to H-CoT and X-Teaming, challenging attack mechanisms that the model wasn't even trained on during the initial training.
Priya: The authors are showing that ReSI largely preserves reasoning, knowledge, instruction-following capabilities, and benign compliance even while improving safety.
Nadia: It’s about using experimental feedback to select and combine alignment methods while assessing safety improvements alongside instruction following and over-refusal for the next model update.
Elias: This whole paper is about building a validated model update as the target for the next round of improvement.
Priya: It supports automated experimentation as a practical approach to learning from newly discovered vulnerabilities in AI systems.
Conclusion: Nadia: We’ve walked through the details of ReSI, which is this recursive safety improvement framework that applies diverse red-teaming methods to identify vulnerabilities in the current target model.
Elias: It functions by developing training recipes through automated research and adopting a validated model update as the target for the next round of improvement.
Priya: Ultimately, this approach strengthens resistance to identified attacks while retaining existing defenses and demonstrates resilience to unseen risks.
Nadia: The authors have shown this framework can achieve resistance across different models and even resilience against unseen attacks on X-Teaming where four ReSI-trained models substantially outperform every frontier model comparator.
Elias: It’s a way to systematically tackle safety by constantly feeding the system new threats and refining its defenses in a bounded, operational sense.
Priya: This framework suggests that effective safety updates need training strategies suited to both the available attack data and the model's existing safety behavior for real-world deployment.
Nadia: ReSI is a recursive safety improvement framework that applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes through automated research, and adopts a validated model update as the target for the next round.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel