RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
cs.AI, cs.CL
Submitted: 2026-08-25
Updated: 2026-08-27
Code: https://github.com/jianghoucheng/RePolicy
Terminology
Sources
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- OpenAI GPT-5 System Card
- Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Agents of Chaos
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
- A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection