Reveree: Diagnosing LLM Reverse-Engineering Agents
cs.CR
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reverse engineering (RE) is critical to security tasks such as malware analysis and vulnerability discovery, and large language model (LLM) agents are increasingly able to perform it autonomously.
Terminology
Abstract
Reverse engineering (RE) is critical to security tasks such as malware analysis and vulnerability discovery, and large language model (LLM) agents are increasingly able to perform it autonomously. Capture-the-flag (CTF) RE challenges have become the standard proxy for measuring this capability, but evaluation rests on a single criterion: whether the agent captures the flag. This solve rate reveals neither where in the RE process an agent fails nor whether a success reflects analysis of the binary or recall of a public solution. In this paper, we propose Reveree, a diagnostic framework that scores an LLM RE agent's trajectory at three tiers: solve rate, milestone progress through an eight-stage RE schema, and a behavioral profile of its actions. Comprehension stages are scored by an outcome-blinded LLM judge validated against a human expert; all other stages are verified deterministically. Using Reveree, we evaluate nine frontier models and four prompting strategies on 88 picoCTF and NYU-CTF challenges. We find that the base model dominates performance, whereas prompting strategy is a secondary, model-dependent effect. Surprisingly, larger, newer, or costlier models are not reliably stronger. We also find that failures concentrate at the comprehension stages of the RE process, and that extra budget, persistence, or reasoning effort rescues few of them, pointing to a competence limit rather than a resource limit. Regarding memorization, while models reproduce picoCTF flags from challenge descriptions alone, NYU-CTF shows minimal measurable recall, and most solves survive surface perturbation, indicating that genuine analysis coexists with memorization. We release Reveree to the community.
Sources
- AutoPenBench: Benchmarking Generative Agents for Penetration Testing
- Chain-of-Thought Prompting of Large Language Models for Discovering and Fixing Software Vulnerabilities
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- GPT-4 Technical Report
- HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
- CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
- AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs