CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance
cs.CR, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/apache/caldera
Project page: https://cyberclear-bench.github.io/cyberclear
Terminology
Sources
- AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- AgentScope: A Flexible yet Robust Multi-Agent Platform
- AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- GPT-4o System Card
- MITRE ATT&CK Applications in Cybersecurity and The Way Forward
- CAM-LDS: Cyber Attack Manifestations for Automatic Interpretation of System Logs and Security Alerts
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- NODLINK: An Online System for Fine-Grained APT Attack Detection and Investigation
- Kitsune: An Ensemble of Autoencoders for Online Network Intrusion Detection
- LLM-driven Provenance Forensics for Threat Investigation and Detection
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- ProvAgent: Threat Detection Based on Identity-Behavior Binding and Multi-Agent Collaborative Attack Investigation
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs