How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
cs.CR, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- LLM Agents can Autonomously Hack Websites
- Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
- CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
- CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
- Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
- An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
- Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
- D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs