Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
Gal Engelberg, Michael Arenzon, Leon Goldberg
cs.CR
Submitted: 2026-07-29
Comments: 14 pages, 3 tables. Code: https://github.com/OpenSecurityAI/osb ; datasets: https://huggingface.co/OpenSecurityAI
Code: https://github.com/OpenSecurityAI/osb
License: http://creativecommons.org/licenses/by/4.0/
The gist: Enterprises are moving toward autonomous cyber defense: agentic AI that builds situational awareness of an organization's security state and reasons from it to assessments, decisions, and actions.
Terminology
Abstract
Enterprises are moving toward autonomous cyber defense: agentic AI that builds situational awareness of an organization's security state and reasons from it to assessments, decisions, and actions. This rests on a holistic view of the enterprise's security state, the continuous, cross-vendor picture of identities, cloud and infrastructure, data, applications, and their configurations that security posture management assembles. As agents take on this work, what matters is not whether an agent can produce an answer but whether it should be trusted to. The field cannot yet answer this question. Real enterprise environments are private, cross-vendor, and deeply correlated, and none is exposed publicly as a shared, queryable target for evaluating such agents end to end. We call this the environment data gap. We present Open Security Benchmark (OSB), a framework that benchmarks agentic AI on this work. OSB surfaces a curated enterprise environment - a frozen, holistic view of the security state - and evaluates posture investigation across two modalities: text-to-SQL over a relational snapshot and each vendor's native API over a served instance of the same environment. Freezing the environment pins the target state as an immutable snapshot and anchors answers to a closed-form ground truth. OSB is built from five components: a data layer, a task and evaluation-set layer, a multi-dimensional scoring layer, a minimal auditable harness, and a bring-your-own path that serves public comparison and private tenant evaluation from one substrate. We instantiate the framework with two identity-security packs and a family of synthetic-organization environment datasets spanning multiple scales, and chart its extension to further posture subdomains, investigation modalities, and defense stages from assessment toward remediation.
Sources
- Automated Cyber Defence: A Review
- CybORG: A Gym for the Development of Autonomous Cyber Agents
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
- Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility
- Cross-Vendor Sola ISPM Benchmark: Evaluating Agentic AI for Federated Identity Security Reasoning
- CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
- CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
- CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- Gorilla: Large Language Model Connected with Massive APIs
- Large Language Models are not Fair Evaluators
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Training language models to follow instructions with human feedback
- Deep reinforcement learning from human preferences
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs