SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
cs.CR, cs.AI, cs.CL
Submitted: 2026-07-29
Code: https://github.com/Alibaba-NLP/qqr
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
- Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- IRCopilot: Automated Incident Response with Large Language Models
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- CocoaBench: Evaluating Unified Digital Agents in the Wild
- CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs