REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming
Nicolas Koller, Andreas U. Schmidt
cs.SE, cs.AI, cs.CR, cs.PL
Submitted: 2026-07-07
Comments: 10 pages, 5 figures; accepted for publication to the 23rd International Conference on Applied Computing 2026, Lisbon October 24-26,2026
Code: https://github.com/NicolasKol/reforge
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Binary Code Summarization: Benchmarking ChatGPT/GPT-4 and Other Large Language Models
- Assemblage: Automatic Binary Dataset Construction for Machine Learning
- BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models
- Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
- Symbol Preference Aware Generative Models for Recovering Variable Names from Stripped Binary
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties