Why Does Agentic Safety Fail to Generalize Across Tasks?
cs.LG, stat.ML
Submitted: 2026-05-07
Updated: 2026-09-27
Code: https://github.com/Tomerslortau/agentic-safety-generalization
Terminology
Sources
- GitHub's Copilot Code Review: Can AI Spot Security Flaws Before You Commit?
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- A System for Human-AI collaboration for Online Customer Support
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Safe Exploration in Continuous Action Spaces
- RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning
- WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Generalizing from a few environments in safety-critical reinforcement learning
- A CMDP-within-online framework for Meta-Safe Reinforcement Learning
- Adam: A Method for Stochastic Optimization
- Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed
- In-context Reinforcement Learning with Algorithm Distillation
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Deep Reinforcement Learning: An Overview
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks