APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry
cs.CR, cs.AI
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/darpa-i2o/Transparent-Computing
Terminology
Sources
- Retrieval-Augmented LLMs for Security Incident Analysis
- DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
- An Empirical Study of Observability Limits in Advanced Software Supply Chain Attacks
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs