Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation
cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Project page: https://chaunceykung.github.io/evolutionary-safety-rsi
Terminology
Sources
- Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
- Concrete Problems in AI Safety
- Emergent social conventions and collective bias in LLM populations
- Sabotage Evaluations for Frontier Models
- Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
- The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
- Agent Memory Is a Surface for Endogenous Authorization Laundering
- When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Automated Researchers Can Mitigate Well-characterized Alignment Failures
- TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Systems Security Foundations for Agentic Computing
- J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection