SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research

arXiv:2603.17883 · cs.CR · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research".

Jane: The paper was written by Zehra Karadağ, Simon Klix, René Walendy, Felix Hahn, Kolja Dorschel et al. from Ruhr University Bochum and Max Planck Institute for Security and Privacy.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Discussion of Challenges: Tom: So we've looked at how the paper is structured, but now let's talk about what it has found that’s hindering progress in "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research." The authors have really pinpointed some common issues across all three domains.

Jane: They found that published techniques are often compared very sporadically, which is a huge problem because we can't tell if one method is better than another without direct comparison.

Tom: That lack of comparison is a critical issue; it makes it hard to see where the real advancement in HRE has actually happened.

Lu: The paper emphasizes that prior results are rarely used as direct foundations for subsequent work, which suggests a lot of redundant effort is being repeated across the research community.

Meng: I found that concept frustrating from an engineering standpoint, because we're constantly reinventing the wheel when we should be building on solid prior research findings.

Lalam: It feels like this lack of reusability creates a systemic barrier to genuine cumulative progress in HRE, which is a big deal for the future of security.

Tom: The paper’s analysis shows that the scarcity of high-quality artifacts is a major contributor to this lack of reproducibility and comparability in "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research."

Jane: It's not enough for us to just have a method; we need the actual tools and data that allow us to test it properly.

Lu: The authors identify technical hurdles, like how difficult it is to generalize segmentation models across different technology nodes in IC reverse engineering.

Meng: And even if we get the images, getting them from a real-world advanced chip is quite a barrier because of access issues and specialized equipment costs.

Lalam: This research identifies not just technical problems but structural ones that are impeding our progress as well.

Tom: It’s clear that "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research" is pointing out the deep-seated issues, not just the symptoms.

Jane: We need to understand these challenges before we can move on to how the paper suggests fixing them.

Discussion of Improvements: Tom: Moving into "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research," let's look at the recommendations for improvement. The authors propose three cross-cutting opportunities that cut across technical, organizational, and legal dimensions.

Jane: They are suggesting ways to make HRE more rigorous and scalable by focusing on making the data reusable and standardized in a way that supports community-wide efforts.

Tom: It's not just about better tools; it's about creating a shared environment for the work, which is something really important.

Lu: The paper suggests moving toward an artifact-centric approach to improve reproducibility, meaning the focus should be on providing working implementations rather than just theoretical descriptions.

Meng: I like that idea of engineering work as a first-class output, because having a documented tool that works is far more valuable than just knowing the theory behind it.

Lalam: The idea of standardized benchmarks and evaluation metrics is crucial for ensuring comparability, which moves us away from comparing apples to oranges in different papers.

Tom: This standardization is what makes sense, as it allows us to objectively measure progress across different research paths within "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research.

Jane: They are also advocating for clearer legal frameworks, which addresses the uncertainty around data sharing and collaboration in this field.

Lu: That’s a necessary point; the legal hurdles often act as an invisible barrier that prevents people from sharing their hard work effectively.

Meng: From an implementation standpoint, it's vital to have those shared benchmarks so we can test our new techniques against a real, consistent target.

Lalam: This ensures that the advances we make are meaningful and verifiable by strengthens the trust in the hardware itself.

Tom: It looks like a holistic solution to problems is being put forward in "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research."

Jane: We're definitely seeing clear paths for how this research is meant to be integrated into the future, but we need a final summary of what all this means.

Conclusion: Tom: We have covered so much ground with "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research," from the initial look at the authors and structure, to the challenges and finally, to the suggested improvements.

Jane: It's a massive effort that really lays out the current state of play for anyone working in hardware security or design.

Tom: The main thing I take away is that despite all these technical advances, progress in HRE is still difficult to reproduce because of a lack of shared artifacts—only four percent are actually reproducible.

Lu: That statistic really highlights the problem, making it clear that the theoretical work is not enough; we need real-world validation and practical tools.

Meng: I hope that industry will take these suggestions seriously, especially regarding providing legally cleared reference designs to help accelerate our development cycles.

Lalam: The goal is to move toward a more coordinated and scalable discipline where hardware trust can be verified effectively for the greater good.

Tom: It’s an enormous project, this "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research," because it’s not just one person’s work; it's the collective effort of a huge community.

Jane: We appreciate the thoroughness of this paper and we hope that all those who are interested in hardware security will be able to benefit from this detailed guide.

Lu: I think this work provides a starting point for much more collaborative, systematic research in the future.

Meng: It gives us a clear roadmap for how to build better tools and processes moving forward, which is extremely practical.

Lalam: I believe that by acknowledging these challenges, we are taking a step toward building a more transparent and trustworthy technological culture.

Conclusion: Tom: So, we're wrapping up our discussion of "SoK: From Silicon to Netlist and Beyond - Two Decades of Hardware Reverse Engineering Research" after a really deep dive into the findings. It’s clear that this paper provides a definitive map of where HRE has been and where it is struggling right now.

Jane: It's an incredibly valuable resource because, as we saw, it summarizes two decades of research in IC, FPGA, and netlist RE across all those one hundred eighty-seven publications.

Lu: The scope of this work is impressive; the sheer volume of literature analyzed really shows the breadth of technical knowledge available in this field.

Meng: But I’m glad we focused on how to fix that complexity, because knowing exactly what’s wrong isn' just as important as the solution, right?

Lalam: The ultimate vision is a shift toward a more coordinated and trustworthy HRE community where the trust in silicon is verifiable by design.

Tom: That's exactly what I hope we see—a real commitment to making that kind of open, rigorous approach possible for everyone.

Jane: It’s been really good having these conversations with all of you; it helps put all this technical research into perspective for our listeners.

Lu: I think the idea that' the future needs a better foundation is something we can really build upon creatively in our next projects, too.

Meng: Exactly, and by defining those clear paths forward, we can start thinking about how to actually implement these solutions in real hardware design workflows.

Lalam: Ultimately, for the public good, ensuring that hardware is secure requires this level of transparency and accountability across all industries.

Tom: It’s definitely a roadmap for a truly more reliable digital infrastructure.

Zehra Karadağ, Simon Klix, René Walendy, Felix Hahn, Kolja Dorschel, Julian Speith, Christof Paar, Steffen Becker

Ruhr University Bochum · Max Planck Institute for Security and Privacy

cs.CR

Submitted: 2026-08-19

Updated: 2026-08-20

Code: https://github.com/YosysHQ/prjtrellis

Importance score: 86/100

The gist: The paper presents a systematization of knowledge concerning Hardware Reverse Engineering (HRE), an essential capability for establishing and verifying trust in the hardware supply chain, which

Key concepts

Hardware Reverse Engineering (HRE)
This field involves analyzing hardware, such as chips or netlists, to understand their design and function. The research paper examines two decades of this work, identifying how the industry has progressed and where systemic barriers exist.
Lack of Reproducibility
A major hurdle in HRE is that prior results are rarely used as direct foundations for new work. This leads to redundant effort, as researchers often repeat previous findings instead of building upon them.
Standardized Benchmarks
To ensure comparability across different research paths, the paper suggests using standardized benchmarks and evaluation metrics. This allows researchers to objectively measure progress against a real, consistent target.

Terminology

Summary

The paper presents a systematization of knowledge concerning Hardware Reverse Engineering (HRE), an essential capability for establishing and verifying trust in the hardware supply chain, which supports critical security tasks including vulnerability discovery, design verification through hardware Trojan detection, and counterfeit avoidance.

The study characterizes technical methods across the HRE workflow and identifies technical and organizational challenges that impede research progress by analyzing a corpus of 187 peer-reviewed publications spanning two decades (1985 to 2025). The analysis is structured around three primary research questions:

RQ1: What technical approaches exist for HRE?

The paper details the technical approaches across the three main domains:

  • IC Reverse Engineering: This process involves five stages: I1 Sample Preparation (e.g, depackaging and delayering), I2 Imaging (typically using SEM), I3 Stitching & Stacking, I4 Semantic Segmentation (identifying wires and transistors via classical computer vision or deep learning), and I5 Netlist Extraction, abstracting the physical layout into a network of Boolean logic gates.

  • FPGA Reverse Engineering: This involves four major steps: F1 Bitstream Format RE (understanding the proprietary file structure), F2 Bitstream Extraction (tapping memory buses or dumping memory), F3 Bitstream Decryption (though the study focuses on unencrypted bitstreams), and F4 Bitstream Conversion, converting the bitstream into a gate-level netlist.

  • Netlist Reverse Engineering: This process is broken down into four stages: N1 Partitioning (assigning gates to functional cores or categories), N2 Module Identification (identifying functionality using structural or functional methods), N3 Algorithmic Recovery (recovering high-level abstractions like FSM state graphs), and N4 Sensemaking, which interprets the recovered algorithm in its operational context.

RQ2: What are the technical and organizational challenges hindering the advancement of HRE?

The analysis revealed several substantial overlaps in challenges across HRE subdomains:

  • Industry Gap: A significant challenge is that academic HRE capabilities do not seem to exist for advanced processes, such as FinFET or GAAFET nodes, with publications predominantly considering planar CMOS nodes from 1999 to 2010.

  • Generalizability Issues: Recognition methods struggle to generalize across technology nodes and layers due to domain shift problems, stemming from variations in shapes and materials used in ICs.

  • Error Propagation: The success of extracting a netlist depends on image recognition algorithms; contamination (like dust) or suboptimal sample preparation can lead to irreversible errors, making perfect results difficult to achieve.

RQ3: To what extent do HRE publications provide artifacts that are available, functional, and reproducible?

The evaluation of the artifacts accompanying the 187 papers yielded concerning results:

  • Only 31 publications (17%) claimed to provide publicly available artifacts, corresponding to 30 distinct artifacts in total.

  • Of the 20 artifacts with exercisable tools, only seven allowed us to reproduce at least some of the key results reported in their corresponding publications. This is noted as being exceptionally low (4%).

Conclusion and Recommendations:

The study concludes that while technical advances exist, progress remains difficult to reproduce or build upon. To address these persistent technical and organizational challenges, the authors suggest coordinated action targeting three cross-cutting opportunities:

  1. Improving Reproducibility and Reusability: Promoting artifact-centric practices.

  2. Enabling Rigorous Comparability: Establishing standardized benchmarks and evaluation metrics.

  3. Improving Legal Clarity for Public HRE Research: Addressing constraints imposed by licensing terms, non-disclosure agreements, and export-control regulations that limit data sharing and collaboration.

Improvements for AI systems

The following improvements are based on the systemic failures identified in the systematic review of 187 Hardware Reverse Engineering (HRE) publications. These improvements define a new AI-driven framework, The HRE Knowledge and Validation Engine (HRE-KVE), designed to mitigate fragmentation, enforce reproducibility, and establish standardized benchmarks across IC, FPGA, and Netlist domains.


Improvement: Implementation of a Unified Multi-Modal Knowledge Graph (UMKG) that automatically ingests all 187 publications, mapping the technical methods across IC, FPGA, and Netlist domains into a single relational structure.

  • Mechanism: A specialized Large Language Model (LLM) fine-tuned on the HRE Codebooks (Table 2 & 3) will perform automated entity recognition (NER) to extract:

  • Input/Output Dependencies: Identifying which specific input artifacts (e.g., GDSII, Bitstream, Flattened Netlist) lead to which output results.

  • Causal Linkages: Mapping the sequence of steps (I1 to I2 to I3 to I4 in IC RE) and identifying where published techniques fail or transition between different sub-disciplines.

  • System Capability: The HRE-KVE can generate a dynamic, holistic State of Art visualization that identifies knowledge gaps (e.g., the lack of a unified workflow for stitching and stacking across different IC manufacturing nodes) and instantly cross-reference the optimal techniques for achieving specific high-level goals (e.g., Vulnerability Discovery in GAAFET nodes).

Improvement: Development of an automated, multi-stage Artifact Validation and Verification (AVV) pipeline to address the observed 96% failure rate in reproducing published results.

  • Mechanism: The AVV pipeline uses a standardized Artifact Badge protocol (Availability, Functionality, Reproducibility). When a paper claims to provide artifacts:

  • Automated Functional Testing: A containerized execution environment (Docker/Singularity) validates the tool's executability. It runs automated tests against the claimed inputs (e.g, running a bitstream through the tool or applying SEM image filters).

  • Automated Reproducibility Check: The system executes predefined test cases derived from the paper’s reported key results and calculates a quantifiable error metric (e.g., Intersection over Union (IoU) for segmentation, or percentage match for logic gate equivalence).

  • System Capability: The HRE-KVE will provide a dynamic Reproducibility Score for every published artifact. It flags papers that are overly idealized (assuming error-free inputs) and warns users when the required input binaries or dependency lists are missing, ensuring that researchers only rely on validated tools.

Improvement: Creation of a standardized Benchmark Generation and Comparison Engine (BGC-Engine) to address the lack of comparable targets across publications.

  • Mechanism: The BGC-Engine creates synthetic, but physically accurate, reference datasets (e.g., 800k synthetic SEM images for training data; small, verified netlists for specific logic functions). It then runs all published techniques against these standardized benchmarks using agreed-upon metrics (e.g., ESD or IoU).

  • System Capability: The HRE-KVE allows users to perform Comparative Analysis. Instead of relying on fragmented literature, a user can query the system: Which technique among those targeting Xilinx Virtex-5 and AES S-boxes achieves the highest functional accuracy when tested against the standardized BGC benchmark? This provides objective, quantitative comparisons that were previously impossible.

Improvement: Implementation of a Dynamic Error Propagation Tracer (DEPT) to model and predict failure modes in physical reverse engineering processes (IC/FPGA).

  • Mechanism: DEPT integrates the known vulnerabilities identified in the literature (e.g., suboptimal delayering, dust particle contamination, poor SEM parameter choices). It models how an error introduced at Step I2 (Imaging) propagates through Step I3 (Stitching) and predicts its impact on the final output quality of Stage I5 (Netlist Extraction).

  • System Capability: The HRE-KVE can simulate what-if scenarios for a given physical process. For instance, a user can input parameters for an SEM imaging run and receive a probabilistic assessment of how likely that specific input is to lead to segmentation errors or misclassification in the resulting netlist, enabling proactive mitigation before running expensive physical experiments.

The HRE-KVE transforms the existing fragmented landscape into a unified, actionable knowledge base. It moves beyond merely listing techniques and provides:

  1. Holistic Synthesis: A single view of the entire HRE pipeline, identifying optimal paths across IC/FPGA/Netlist domains.

  2. Trust Quantification: Objective scoring of research quality based on reproducibility and artifact integrity (Trust Score).

  3. Comparative Benchmarking: Quantitative comparison of techniques against standardized, verified targets.

  4. Predictive Modeling: Simulation of physical process errors to minimize downstream technical failures in reverse engineering attempts.

Sources

Related papers