AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination
cs.CR
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: To be published in the 19th ACM Workshop on Artificial Intelligence and Security (AISec 2026) co-located with CCS 2026
Code: https://github.com/Golim/agent-lsd
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- StruQ: Defending Against Prompt Injection with Structured Queries
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- LLM Agents can Autonomously Hack Websites
- Prompt Injection attack against LLM-integrated Applications
- HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
- CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs