HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control
cs.CR, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- Defeating Prompt Injections by Design
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Reliable Weak-to-Strong Monitoring of LLM Agents
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Meta-Harness: End-to-End Optimization of Model Harnesses
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment
- Auditing Agent Harness Safety
- SafeAgent: A Runtime Protection Architecture for Agentic Systems
- Agent Safety Alignment via Reinforcement Learning
- A Framework for Formalizing LLM Agent Security
- Agent Security Needs Redefinition through a Holistic Framework
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Qwen3 Technical Report
- ReAct: Synergizing Reasoning and Acting in Language Models
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs