DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
Xuyang Liu, Yibin Han, Zhenwei Zhang, Kai Chang, Zhiwei Xu, Tian Qiu, Weixian Deng, Jiabao Gao, Xiaolin Peng, Hai Wan, Xibin Zhao
cs.CR, cs.AI
Submitted: 2026-08-04
Code: https://github.com/abrahaamm/DiagChain
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions.
Terminology
Abstract
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions. However, existing benchmarks mainly evaluate final outputs or aggregate accuracy, providing limited insight into how errors arise and propagate across intermediate reasoning stages. We present DiagChain, a diagnostic benchmark for evidence-grounded attack chain reconstruction that enables stage-wise evaluation of LLM agents. DiagChain includes MAIN-69, a suite of 69 scenarios spanning multiple operating systems, evidence noise levels, and chain lengths. It further introduces Evidence-Centric Retrieval-Augmented Generation (ECRAG), which couples evidence retrieval with an evolving structured representation of the reconstructed chain. Five complementary metrics are introduced to assess distinct stages of the reconstruction process and support systematic failure diagnosis. Based on evaluations using 6 LLMs, DiagChain reveals that even the strongest configuration succeeds on only 39.6% of the 849 reference steps in MAIN-69. Our analysis further shows that smaller models struggle with the more basic task of incorporating retrieved evidence into their outputs, whereas larger models can proceed to later steps, where correctly ordering that evidence becomes the main bottleneck. These results validate the importance of diagnostic evaluation beyond end-to-end accuracy and provide actionable insights for improving evidence-grounded cybersecurity agents.
Sources
- Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations
- SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents
- OMNISEC: LLM-Driven Provenance-based Intrusion Detection via Retrieval-Augmented Behavior Prompting
- Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
- Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval
- Before You Hand Over the Wheel: Evaluating LLMs for Security Incident Analysis
- AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
- LLM-driven Provenance Forensics for Threat Investigation and Detection
- HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection
- An Empirical Study of Observability Limits in Advanced Software Supply Chain Attacks
- FuseChain: Runtime Evidence Reconstruction for Software Supply-Chain Attacks
- TGCM: Topic-Guided Consistency Modeling for One-Step Disentanglement of Interleaved APT Technique Sequences
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- Deep Learning-based Intrusion Detection Systems: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs