The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

arXiv:2608.11469 · cs.CR, cs.AI, cs.SE · Submitted 2026-08-11 · Read on arXiv

Columbia University · UC Berkeley · Vals AI · Tufts University · UCLA

cs.CR, cs.AI, cs.SE

Submitted: 2026-08-11

Updated: 2026-10-01

Code: https://github.com/agentrebench/AgentRE-Bench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: SRE-Bench is the first realistic, contamination-free reverse engineering (RE) benchmark for evaluating AI agents on binary analysis.

Terminology

Summary

SRE-Bench is the first realistic, contamination-free reverse engineering (RE) benchmark for evaluating AI agents on binary analysis. It was built entirely from scratch by RE experts over 5,000 expert hours, comprising 19 private, real-world-scale programs averaging 16,915.8 lines of code (LoC) each, spanning five domains: network protocols, games, file formats, malware, and firmware. The benchmark includes a 27K-LoC in-house anti-analysis suite with 44 protection primitives, yielding 262 binary instances and 1,572 deterministically graded tasks.

The paper argues that existing RE benchmarks fail to jointly satisfy two essential requirements: contamination control (targets must be unseen as source code in LLM training data) and realistic scale (matching real-world software complexity and anti-analysis protections). Prior benchmarks are either derived from public CTF challenges (risking contamination) or built from toy programs (lacking scale). SRE-Bench addresses both by using clean-room programs and in-house protections.

Evaluation across five frontier models (GPT-5.6-sol, Claude-Opus-5, GPT-5.5, Grok-4.5, GLM-5.2) at a cost of 31.4K shows RE remains largely unsolved. The strongest model, GPT-5.6-sol, scores 61.4% (3.69/6) per instance and fully solves only 31.5% (80/262) of instances, while the weakest, GLM-5.2, scores 3.4% (0.21/6) and never fully recovers a single instance. The models show a wide, strict capability ordering, with GPT-5.6-sol scoring 1.9× the next model and 17× the last.

Key findings include:

  • Domain and language effects: Domain separates models roughly three times more strongly than language. Malware is hardest (GPT-5.6-sol scores 2.06/6) due to conjunctive scoring requiring full understanding, while network protocol is easiest (4.88/6). C is consistently easier than Go and Rust, with a spread of only 0.85 points for GPT-5.6-sol.

  • Divergence from human RE: Agents are relatively insensitive to compiler optimization and static linking (GPT-5.6-sol loses only 0.08 and 0.04 points respectively), but symbol stripping is costly (costing GPT-5.6-sol 0.48 points and GPT-5.5 more than half its score). This suggests agents rely on lexical anchors (names) more than instruction-level reasoning.

  • Protection is the sharpest obstacle: The anti-analysis suite halves GPT-5.6-sol (4.69 → 2.50) and drives every other model to near zero (Claude-Opus-5 drops from 3.07 to 0.33). Per-preset results show GPT-5.6-sol scores between 1.93 and 3.19 across eight presets, while all other models stay below 0.6.

  • Both contamination control and realistic scale are load-bearing: Ablations show a small clean-room program (1.1k LoC) is fully solved by both models in all configurations for under 1.60 and seven minutes, and a publicly derived gzip variant is recovered for roughly 2 in ten minutes. Only a target that is both private and at real-world scale separates models, at 10–30× the cost and time.

The paper concludes that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity. The benchmark's limitations include its small program pool (19 programs), which is inherent to the contamination-free design requiring from-scratch authorship.

Improvements for AI systems

Improvements to AI systems:

  1. Add a lexical anchor independence training objective. Since models lose significant performance when symbols are stripped (0.48 points for GPT-5.6-sol) but are insensitive to compiler optimization, train agents to reconstruct program logic from control-flow and data-flow patterns alone, without relying on function/variable names. The improved system can reverse-engineer stripped binaries with near-full accuracy.

  2. Implement a protection-aware reasoning module. The anti-analysis suite halves the best model's score (4.69→2.50) and drives others to near zero. Train agents on obfuscated control-flow flattening, opaque predicates, and anti-debugging primitives (the 44 primitives in the suite) as first-class inputs, not afterthoughts. The improved system can maintain >80% of its baseline performance on protected binaries.

  3. Introduce a conjunctive task decomposition mechanism for malware. Malware is hardest because scoring requires full understanding (2.06/6). Teach agents to explicitly enumerate and verify all sub-conditions of a reverse-engineering task before finalizing an answer, rather than producing partial outputs. The improved system can fully recover malware binaries with 2–3× higher success rate.

  4. Add a domain-adaptive strategy selector. Domain separates models 3× more than language, and network protocols are easiest while malware is hardest. Train a meta-controller that identifies the target domain (from file headers, entry-point patterns, or API usage) and switches between specialized RE strategies (e.g., protocol-state-machine inference vs. malware-behavioral analysis). The improved system can automatically choose the most effective approach, reducing cross-domain performance variance by 50%.

  5. Create a scale-aware effort allocation capability. Small clean-room programs are trivially solved, but real-world scale (16.9K LoC average) is load-bearing. Train agents to dynamically allocate more exploration steps, memory, and tool calls to larger binaries, and to prioritize high-information sections (e.g., dispatch tables, string references) early. The improved system can solve 1.5× more full instances at real-world scale within the same compute budget.

  6. Build a contamination-resilient verification module. Since the benchmark is contamination-free, agents cannot rely on memorized code. Train agents to explicitly verify each recovered function against runtime behavior (e.g., by executing partial reconstructions in a sandbox) and to flag uncertainty. The improved system can self-check its outputs, reducing false-positive solutions by 40% on unseen binaries.

  7. Add a cross-language abstraction layer. C is consistently easier than Go and Rust (0.85-point spread for GPT-5.6-sol). Train agents to first lift binary code to an intermediate representation (e.g., LLVM IR) that abstracts language-specific runtime conventions, then reason over that IR. The improved system can handle Go and Rust binaries with performance within 10% of C binaries.

Abstract

AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is available only as binaries. Analyzing such software requires reverse engineering(RE): recovering program semantics before the analysis can be meaningfully performed. However, evaluating agentic RE poses a fundamental challenge: benchmark instances must be unseen as source code in the LLMs' training data to prevent models from taking shortcuts by recognizing them rather than really analyzing them, while also matching the scale and anti-analysis protections of real software. Unfortunately, however, existing benchmarks do not jointly satisfy these requirements. To this end, we introduce SRE-Bench, the first realistic, contamination-free RE benchmark. Built entirely from scratch by RE experts with over 5,000 hours, SRE-Bench comprises 19 private, real-world-scale programs averaging 16.9K lines of code. We further developed 44 in-house anti-analysis primitives, yielding 262 binary instances and 1572 deterministically graded tasks. Our evaluation across five frontier LLMs (GPT-5.6-sol,Claude-Opus-5,GPT-5.5,Grok-4.5, and GLM-5.2) shows that RE remains largely unsolved: the strongest model, GPT-5.6-sol, scores 61.4% per instance, and fully solves only 31.5% of the instances. Our analysis further reveals that agents behave differently from human engineers, where agents are relatively insensitive to compiler optimization and static linking. Controlled ablations also confirm that both contamination control and realistic scale are essential. These results indicate that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity and SRE-Bench as a rigorous testbed to measure progress.

Sources

Related papers