Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations
cs.CR, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-28
Code: https://github.com/openai/codex
Terminology
Sources
- Cross-Vendor Sola ISPM Benchmark: Evaluating Agentic AI for Federated Identity Security Reasoning
- Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
- Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
- WideSearch: Benchmarking Agentic Broad Info-Seeking
- Why Do Multi-Agent LLM Systems Fail?
- SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents
- Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
- Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations
- DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
- Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
- Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
- TAC: Hybrid IAM Privilege Escalation Detection
- Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Over-Searching in Search-Augmented Large Language Models
- DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs