EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild

arXiv:2604.01554 · cs.CR, cs.LG, cs.SE · Submitted 2026-04-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild".

Elias: Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: Welcome back everyone; today we’re diving into this paper called "EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild." This sounds like it tackles a really hard problem in software security, which is Binary Function Similarity Detection, or BFSD.

Elias: I agree, Nadia; it seems like the core issue they are addressing is that current models struggle to generalize because their training data doesn't cover enough real-world variations. This paper introduces EXHIB as a way to fix that by creating a comprehensive set of benchmarks.

Priya: From my perspective, what I’m interested in is exactly what kind of diversity they’re talking about when they collect these datasets, and how that affects the actual data quality for privacy and measurement purposes.

Nadia: Exactly, Priya; the paper frames this as a lack of a universal benchmark, suggesting that existing datasets are too narrow in scope and often only focus on a small set of transformations or binary types. EXHIB aims to fix this by covering three levels: low-level differences, mid-level differences through obfuscation, and high-level semantic differences.

Elias: That’s the crucial part; it means they aren't just looking at minor recompilation settings anymore; they’re explicitly trying to test how similarity detection holds up when the underlying structure changes in ways that matter semantically.

Nadia: Right, and they built this benchmark using five realistic datasets collected from the wild, which include standard projects, firmware from various vendors, malware samples, obfuscated code with different tools like Obfuscator-LLVM or Tigress, and even a semantic dataset where multiple participants write solutions to the same problem.

Priya: The inclusion of those diverse sources is interesting; it suggests they are trying to capture variations that aren't just about compiler flags but about how a function is actually being used or implemented differently in practice.

Elias: And their evaluation setup is what makes this paper stand out, as they test nine representative models across three major approaches: fuzzy hashing, graph-based learning, and NLP-inspired modeling to see where each paradigm performs best.

Title and authors: Nadia: It’s a pretty thorough comparison; they show that the performance of these models isn't uniform across all types of binary variations; actually, they found performance degradations of up to thirty percent on firmware and semantic datasets compared to standard settings <ref:2604.01554#pg0,performance degradations of up to 30% on firmware and semantic datasets>.

Priya: That thirty percent drop is significant, especially when it happens on those high-level semantic differences; it shows that robustness isn't a single property you can just claim for any model applied universally <ref:2604.01554#pg0>.

Elias: And the paper points out a key limitation of prior benchmarks: most focus primarily on low-level recompilation settings, while mid- and high-level variations remain underexplored, which is why they felt the need to construct EXHIB to systematically cover all three categories.

Nadia: So, in summary, this paper introduces EXHIB as a way to move beyond narrow testing by systematically covering low-, mid-, and high-level differences using five distinct real-world datasets.

Priya: What I find compelling is their conclusion that robustness is strongly taxonomy-dependent rather than a uniform property of a given architecture, which really shifts how we think about model reliability in this area.

Elias: And that leads directly into the suggested improvements; the paper doesn't just present data; it suggests ways to make future models better by focusing on where they fail based on those taxonomy results.

Nadia: They suggest integrating a hybrid representation, like combining structural control flow graphs with semantic embeddings, to build a more robust Semantic-Oriented Graph representation that handles source code rephrasing better than current methods.

Priya: That sounds promising because it addresses the gap where static structural similarity models struggle with high-level semantic differences and could potentially close that thirty percent performance gap they observed on those sets <ref:2604.01554#pg0>.

Elias: I also think the evaluation methodology itself needs a change; instead of a single score, they suggest implementing a multi-dimensional performance decomposition framework to explicitly measure robustness against low, mid, and high levels of variation separately.

Nadia: That would give us much better diagnostic tools to understand exactly which type of binary difference is causing a model to fail so we can guide targeted improvements instead of just retraining everything.

Title and authors: Priya: And on the data side, they strongly recommend expanding dataset diversity by prioritizing semantically diverse implementations and incorporating state-of-the-art obfuscation techniques like virtualization into the Obfuscated dataset.

Elias: That’s a necessary step because, as we saw with graph methods excelling on obfuscation, if we want models to be truly reliable for real-world security tasks, they need to handle those complex transformations better.

Nadia: Efficiency is another big concern; the paper notes that while graph methods like HermesSim are accurate, they can be computationally expensive, and there’s a push to develop lightweight architectures that keep that structural modeling power but at a faster speed.

Priya: That efficiency point is vital because if we want these tools to be practical for production environments, we need to balance high accuracy with low latency for things like real-time malware screening.

Elias: I also see the importance of integrating dynamic features; since some models show success by incorporating dynamic micro-traces, developing a pipeline that automatically extracts those traces and feeds them back into the model could help bridge the gap between static structure and runtime behavior.

Nadia: So, it seems like the main thrust here is moving from simply measuring similarity to understanding *why* a model succeeds or fails based on the level of variation it encounters.

Priya: I think that focus on taxonomy-dependent robustness is what will actually help us build security tools that are reliable in unpredictable environments.

Elias: And if we look at the overall picture, the paper "EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild" provides a much clearer map of where current BFSD research needs to focus its efforts moving forward.

Nadia: That’s right, it gives us a solid framework for what makes a similarity detection system truly useful in practice across different security domains.

Priya: It's certainly an important piece of work because it forces the community to look beyond just standard compilation variations and consider the semantic complexity of binary functions.

Elias: Exactly; we have to be careful about what we assume when we build these models, and this paper shows us exactly where those assumptions break down.

The paper's summary: Nadia: So, to recap this paper, it's basically introducing EXHIB as this new way to test how well we can detect if two binary functions are similar when they are actually quite different in real-world ways.

Elias: Exactly, and the big takeaway is that current models aren't robust across all kinds of variations because they only see a narrow slice of reality.

Priya: From my side, what really stands out is that this benchmark forces us to look at the gap between low-level changes and high-level semantic changes in a way we haven't before.

Nadia: Right, and they show performance dips up to thirty percent when you move into testing those high-level semantic differences, which is a pretty big deal for real security analysis.

Elias: That gap suggests that our current similarity detection methods aren't just failing on simple syntax changes but are genuinely missing the underlying meaning of the code.

Priya: And it points us toward the idea that we need to be much more precise in how we measure robustness, moving away from a single score to understanding where a model breaks down.

Nadia: Precisely, and this means if we want to build better security tools, like for vulnerability analysis or malware classification, we need models that handle these diverse variations consistently.

Elias: That consistency is the challenge; what I'm thinking is that the paper’s focus on graph-based methods actually seems like a good direction because they handle those structural relationships better than some of the NLP approaches.

Priya: I agree, and it makes me wonder how this applies beyond just function similarity, like when we think about identifying malicious code across different compilers or hardware platforms.

Nadia: That’s what I want to ask next; if we can build models that handle this level of complexity better, what does that mean for the cost of exploitation? Can we find a way to use this knowledge to make finding vulnerabilities cheaper or easier?

The paper's improvements: Tom: So, we're looking at how they plan to make this work better because right now, the paper highlights that robustness depends entirely on which variation you are testing for.

Nadia: Exactly; they suggest we stop treating similarity detection as a single measure and instead develop a multi-dimensional performance decomposition framework to separate low-, mid-, and high-level variation results.

Elias: That would be useful because it lets us pinpoint exactly where a model is weak, whether it's struggling with compiler differences or with actual source code rephrasing.

Priya: I think that diagnostic capability is what makes this paper so valuable for measurement research, because we can finally see which parts of the system are failing and why.

Nadia: And from an applied security standpoint, if we know a model fails specifically on high-level semantic differences, we can focus our efforts on training it to better understand source code intent rather than just matching syntax.

Elias: That ties back to the representation issue; they suggest integrating hybrid representations that combine structural control flow information with more nuanced token embeddings to handle those semantic leaps.

Priya: It sounds like they are trying to build a system that can bridge the gap between static structure and dynamic meaning, which is a big step for real-world reliability.

Nadia: And I'm curious about the efficiency side; since graph methods are accurate but slow, they’re proposing lightweight graph neural networks or distillation techniques to keep performance high without slowing things down too much.

Elias: That makes sense; we need accuracy that matches those advanced models but inference times that fit into a practical security pipeline for things like real-time threat detection.

Priya: Also, they emphasize expanding the dataset diversity, specifically pushing for more semantic variations and including state-of-the-art obfuscation techniques to ensure the results aren't just skewed toward "standard" scenarios.

Nadia: That’s vital because if we only test on standard stuff, our tools will be useless when dealing with novel malware or new programming language quirks.

Elias: The authors also point out that integrating dynamic feature extraction, like execution micro-traces, could further improve the model's ability to recognize semantic equivalence across different implementations.

Priya: That dynamic input could provide the necessary context to make those structural models more accurate when source code is heavily transformed.

Nadia: So, we’re moving toward a system that’s not just a static checker but one that can adapt its understanding based on the level of complexity it's facing, which is really something to look forward to.

Conclusion: Nadia: So, to wrap things up, this paper on "EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild" essentially shows us that our tools for comparing binary functions are too narrow because they don't account for all the ways code can change in the real world.

Elias: That’s right; it proves that relying on a single test isn't enough when you’re dealing with software coming from diverse sources like firmware or complex source code.

Priya: I think what this means for measurement research is that we need to treat robustness not as an absolute thing, but as something deeply dependent on the specific context of the variation being tested.

Nadia: And for applied security, it suggests that if we build models capable of handling these high-level semantic jumps better, we could potentially make finding vulnerabilities much more accessible and cost-effective across different systems.

Elias: I agree; if the AI can reliably map semantic equivalence even when the syntax is totally different, then identifying functional similarities becomes a much stronger tool for cryptanalysis and security auditing.

Priya: From a privacy standpoint, this means that any model we deploy needs to be rigorously tested against these diverse inputs because those datasets are collected from publicly accessible sources.

Nadia: Exactly; we're dealing with real-world binaries here, not just clean lab examples, so the reliability of the AI has to be proven across a huge spectrum of noise and variation.

Elias: So, this work on EXHIB gives us a much better map for where we need to improve our understanding of binary structure and semantic modeling.

Priya: It really highlights that expanding the variety in data collection is not just about getting more data, but about ensuring that the data reflects the actual complexity found in deployed software.

Nadia: Definitely; this benchmark sets a high bar for what a truly versatile similarity detection system needs to achieve if it wants to be useful in security applications.

Elias: It’s exciting because it points us toward graph-based and hybrid approaches as the most promising path forward, given their ability to model those relationships across different levels of variation.

Priya: I hope future research continues this trend of expanding semantic variability and incorporating more advanced obfuscation techniques into these benchmarks.

Nadia: We've got a lot to think about with this paper on EXHIB; it’s definitely going to influence how we design the next generation of binary analysis tools.

The Ohio State University

cs.CR, cs.LG, cs.SE

Submitted: 2026-04-02

Updated: 2026-10-05

Comments: 13 pages, 7 figures. This is a technical report for the EXHIB benchmark. Code and data are available at https://github.com/fan1192/EXHIB and https://doi.org/10.5281/zenodo.18488936

Code: https://github.com/fan1192/bfsd-a

Project page: https://networkx.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance.

Key concepts

Binary Function Similarity Detection (BFSD)
This is the core task of comparing two computer programs (binary functions) and calculating a score indicating how similar they are. This is crucial for security tasks like identifying malware or determining if one piece of code is a modified version of another.
EXHIB Benchmark
EXHIB is a new set of five realistic datasets designed to test BFSD across different types of binary variations. It covers standard software, firmware, malware, obfuscated code, and semantically equivalent but syntactically different functions to provide a more complete evaluation framework.
Taxonomy of Variation
The paper organizes binary function differences into three levels: low-level (like compiler changes), mid-level (like code obfuscation), and high-level (where the code does the same thing but looks very different). This structure helps researchers systematically test how well models handle different types of similarity challenges.
Graph-Based Learning
This is a machine learning approach where binary functions are represented as nodes in a graph, and connections between them represent structural relationships. The paper finds this method is particularly strong because it can effectively model the underlying semantic structure of the code, making it robust against complex transformations.

Terminology

Summary

Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance.

The gist: Robustness to low- and mid-level binary variations does not generalize to high-level semantic differences.

Introduction and Problem Framing

The paper addresses the lack of a comprehensive universal benchmark for Binary Function Similarity Detection (BFSD), which is framed as taking two binary functions as input and generating a similarity score. Existing datasets are limited in scope, often focusing on a narrow set of transformations or types of binary, failing to reflect the full diversity of real-world applications. Researchers struggle to compare different models effectively due to inconsistent parameterization across frameworks and biases introduced by sampling only typical programs. The authors introduce EXHIB, a benchmark comprising five realistic datasets collected from the wild, designed to systematically cover all variations in binary function differences using a three-level taxonomy: low-level differences (architecture, compilers), mid-level differences (obfuscation transformations), and high-level differences (semantically equivalent but syntactically distinct functions).

Dataset Construction and Taxonomy

EXHIB is constructed by grounding dataset design in the proposed three-level taxonomy to systematically cover all three categories of variation. The five datasets include:

  1. Standard Dataset: Comprises popular open-source projects, where differences are generated by architecture, bitness, compiler, compiler version, and optimization level.

  2. Firmware Dataset: Generated with several hardware vendors and their IDEs; differences are created using 15 versions of ARM compilers and 6 different optimization levels.

  3. Malware Dataset: Source code comes from public repositories; differences are generated by compiling source codes using multiple compilers, compiler versions, and optimization levels.

  4. Obfuscation Dataset: Built from open-source projects where code obfuscation is the only variable used to create differences, employing three obfuscators (Obfuscator-LLVM, Hikari, and Tigress).

  5. Semantic Dataset: Targets high-level differences by collecting source code from programming contest platforms where multiple participants independently implement solutions to the same problem specification.

Model Evaluation and Performance Trends

The authors evaluate 9 representative models spanning multiple BFSD paradigms on EXHIB. The results reveal performance degradations of up to 30% on firmware and semantic datasets compared to standard settings, underscoring substantial generalization gaps. Key trends observed across the datasets include:

(a) Standard Dataset:

(b) Firmware Dataset:

(c) Malware Dataset:

The analysis shows that robustness to low- and mid-level binary variations does not generalize to high-level semantic differences. Specifically, AUC performances are generally 10 − 20% lower on the firmware dataset compared to the standard dataset, and on the semantic dataset, models experience a more pronounced decline, with AUC performances typically 20 − 30% lower than on the standard dataset.

Model Comparison by Paradigm

The study compares models across three main trends: fuzzy hashing, graph-based learning, and NLP-inspired modeling.

(a) Fuzzy Hashing:

FCatalog consistently outperforms FunctionSimSearch across all five datasets, demonstrating that well-designed fuzzy hashing techniques remain competitive. It is noted that FCatalog requires only a single feature extraction pass, making it highly efficient.

(b) Graph-Based Methods:

Graph representations excel at modeling relationships, with Gemini and HermesSim consistently outperforming other categories. HermesSim achieves a perfect score of 1.00 on the Obfuscation Dataset, demonstrating that its Semantics-Oriented Graph (SOG) representation is especially robust to heavy code transformations.

(c) NLP-Based Methods:

NLP-based approaches remain interpretable but are generally outperformed by graph-based methods, especially under heavy obfuscation or cross-architecture/compiler variation. Trex and Asm2Vec perform notably worse, with Asm2Vec recording the lowest AUC, Recall, and MRR@10 across most datasets.

Conclusion and Future Directions

The study concludes that BFSD robustness is strongly taxonomy-dependent rather than a uniform property of a given architecture. Graph-based methods are identified as the most promising for achieving robust performance due to their ability to model semantic relationships and structural patterns within binary functions. The paper stresses the necessity of expanding dataset diversity by including increased semantic variability, expanded obfuscation coverage, and diverse binary formats and languages to enhance the reliability of BFSD models. Furthermore, efficiency metrics, such as feature extraction and inference time, are highlighted as critical criteria for practical model assessment.

Ethics Considerations

The research involved the construction of datasets from publicly accessible sources (open-source projects, vendor firmware images, public malware repositories) and programming contest submissions.

Improvements for AI systems

Based on the provided research paper, here are specific improvements that can be made to existing AI/ML systems in Binary Function Similarity Detection (BFSD), along with what those improved systems could achieve:


)AI System Improvements for BFSD

  1. Improvements in Model Architecture and Representation:

The current models suffer from a lack of robustness across different levels of binary variation (low-level vs. high-level semantic). Specifically, the paper suggests that purely static or token-based embeddings (like Asm2Vec) struggle with high-level semantic differences, while graph representations (HermesSim/Gemini) excel under obfuscation but degrade on semantically diverse source implementations.

[Specific Improvement] Integrate a hybrid representation that explicitly combines structural information from Control Flow Graphs (CFGs) with semantically-aware token embeddings or micro-traces. This could involve developing a Semantic-Oriented Graph (SOG) representation that is more robust to high-level source code variations, potentially by incorporating dynamic analysis features like instruction execution traces into the graph structure.

[Potential Achievement] The improved system would achieve superior generalization across diverse datasets (e.g., Semantic Dataset), maintaining high accuracy when source code changes substantially but retaining the ability to detect similarity even when syntactic structures are completely different, effectively closing the 30% performance gap observed in semantic evaluation.

  1. Improvements in Evaluation Methodology:

The paper demonstrates that performance degradation is taxonomy-dependent, meaning robustness is not a monolithic property. Current metrics often treat models as robust or not, without attributing the failure to specific levels of variation (low, mid, high).

[Specific Improvement] Implement a rigorous, multi-dimensional performance decomposition framework. Instead of reporting a single AUC score on a combined benchmark (like EXHIB), systems should be evaluated on separate metrics for each level: AUC under Low-Level variation (XA/XB), Mid-Level variation (Obfuscated), and High-Level variation (Semantic).

[Potential Achievement] This would allow researchers to pinpoint the exact blind spot of a model. For instance, a system could be identified as robust against compiler changes but highly vulnerable to source code rephrasing, guiding targeted architectural improvements rather than broad retraining.

  1. Improvements in Dataset Curation and Diversity:

The existing datasets are too narrow, focusing heavily on standard compilation variations or single types of obfuscation/source domains (e.g., only one malware family).

[Specific Improvement] Systematically expand the benchmark suite by prioritizing the collection of semantically diverse implementations from programming contests and incorporating state-of-the-art obfuscation techniques (including virtualization) into the Obfuscated dataset. Furthermore, ensure that firmware/malware datasets include a broader spectrum of architectural and compiler combinations as variables.

[Potential Achievement] An AI system trained on this hyper-diverse data will exhibit vastly improved cross-platform and cross-compiler capabilities, making it genuinely useful for real-world software security tasks where binaries originate from unpredictable sources (e.g., diverse IoT device firmware or novel malware).

  1. Improvements in Computational Efficiency:

Graph-based methods (HermesSim) achieve the highest accuracy but incur significantly higher computational costs (hundreds of seconds per 100 functions) compared to fuzzy hashing methods like FCatalog.

[Specific Improvement] Develop lightweight graph neural network architectures or efficient feature extraction pipelines that can approximate the structural modeling power of full GNNs while maintaining the low-latency performance of classical methods. Explore techniques like graph sparsification or distillation from large GNN models into smaller, faster inference networks.

[Potential Achievement] The improved system would achieve a best-of-both-worlds profile: achieving accuracy close to state-of-the-art graph methods while maintaining the speed necessary for practical, real-time vulnerability analysis or malware screening in production environments.

  1. Improvements in Dynamic Feature Integration:

The paper shows that Trex's success on the Semantic Dataset stems from incorporating dynamic micro-traces, which capture runtime behavior consistent across different source implementations—a feature missing from HermesSim’s static SOG approach.

[Specific Improvement] Develop a unified pipeline that automatically instruments binaries to extract representative execution micro-traces during testing. These traces should then be used as additional input features for the similarity model (e.g., augmenting the graph nodes or token embeddings).

[Potential Achievement] The resulting AI system would gain the ability to bridge the gap between static structural similarity and dynamic semantic equivalence, enabling it to correctly classify semantically equivalent but syntactically distinct functions with high precision, even under extreme source-level variation.

Related papers