Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations
cs.CR, cs.CL
Submitted: 2026-06-09
Updated: 2026-09-26
Code: https://github.com/FiveDirections/OpTC-data
Terminology
Sources
- LLM-based event log analysis techniques: A survey
- LogLLM: Log-based Anomaly Detection Using Large Language Models
- Does Prompt Formatting Have Any Impact on LLM Performance?
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- RedChronos: A Large Language Model-Based Log Analysis System for Insider Threat Detection in Enterprises
- Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
- Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection
- BaxBench: Can LLMs Generate Correct and Secure Backends?
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs