Reassembling Distributed Risk: Trajectory-Conditioned Action Generation for Multi-Turn Agent Safety
cs.CR
Submitted: 2026-08-26
Updated: 2026-08-26
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
- LlamaFirewall: An open source guardrail system for building secure AI agents
- Defeating Prompt Injections by Design
- GPT-4 Technical Report
- TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- Ministral 3
- LoRA: Low-Rank Adaptation of Large Language Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- STAC: When Innocent Tools Form Dangerous Chains for LLM Agents
- A Framework for Formalizing LLM Agent Security
- Gemma 4 Technical Report
- Kimi K2.5: Visual Agentic Intelligence
- Progent: Securing AI Agents with Privilege Control
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
- AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs