What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
cs.CR, cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/protectai/llm-guard
Terminology
Sources
- Phi-4 Technical Report
- Quantifying Attention Flow in Transformers
- Qwen Technical Report
- Learning to Attribute with Attention
- ContextCite: Attributing Model Generation to Context
- Defeating Prompt Injections by Design
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- Mistral 7B
- Attention Flows for General Transformers
- Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression
- Progent: Securing AI Agents with Privilege Control
- Gemma: Open Models Based on Gemini Research and Technology
- FCOS: Fully Convolutional One-Stage Object Detection
- LLaMA: Open and Efficient Foundation Language Models
- Attention Is All You Need
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Defending Against Prompt Injection with DataFilter
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs