SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment
cs.AI, cs.CR
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/MaoPopovich/SafeEvolve
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
- EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning
- Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
- ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
- RvB: Automating AI System Hardening via Iterative Red-Blue Games
- Recursive Harness Self-Improvement
- SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Auditing Agent Harness Safety
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- Self-Distilled Agentic Reinforcement Learning
- Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions
- EvoHarness-RL: Learning Runtime Harness Coordination for Self-Evolving Agents
- SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection