SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
summary
The gist
The gist The paper presents a systematization of knowledge (SoK) regarding failure modes in Common Criteria product evaluation, drawing on recurring cross-vendor failures across nine lifecycle
In short
The paper systematizes knowledge about failures during Common Criteria product evaluation by creating a taxonomy across nine lifecycle classes. It identifies recurring cross-vendor issues and proposes a 'design-for-evaluability' framework to guide product teams in preventing these failures before testing begins.
Key concepts
- Lifecycle Taxonomy
- A structured classification system organizing recurring failure modes into nine distinct classes that map the entire product lifecycle, from initial security target scoping through post-certification configuration drift. Failures are only included if they recur across vendors and can be made actionable.
- Design-for-Evaluability Framework
- A set of practices product teams should adopt before evaluation starts. This framework focuses on ensuring the device is drivable as a client and that failures are easily inducible, such as providing clear error logs when a device rejects input.
- Evidence Anchors by Failure Class
- A detailed mapping showing which specific work units (like ASE or ATE) anchor particular failure modes. This links abstract problems to concrete engineering practices and identifies exactly where in the evaluation process a failure occurs, such as boundary redrawing or documentation errors.
Terminology used across episodes
This episode discusses
- SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance · Paper Radio
The paper
SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance · Read on arXiv
Punit Suketu Patel
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance".
Elias: The gist The paper presents a systematization of knowledge (SoK) regarding failure modes in Common Criteria product evaluation,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So, looking at this paper again, "SoK: Failure Modes in Common Criteria Product Evaluation—A Taxonomy and Design-for-Evaluability Guidance," the point is that certification isn't just some checkbox thing anymore; it’s a process full of hidden pitfalls.
Elias: The authors are taking that gap between what a product claims it is and what the evaluation actually finds, and they're trying to close it by creating this structured taxonomy of failure modes. It’s about making the evaluation process itself more predictable.
Priya: What I find interesting is how they frame this entire thing around a design-for-evaluability framework, which means shifting the focus from just fixing problems after testing to preventing those problems in the very early stages of product development.
Nadia: Right, so instead of waiting for a failure to happen during evaluation, the paper suggests we build practices into the design phase—things like making sure you’re drawing your security boundary based on actual evidence rather than just what you *hope* it's going to be.
Elias: That ties into how they structure it: they organize failures across nine lifecycle classes, which follow the path an evaluation actually takes, from scoping the Security Target right up to what happens after you get that certificate.
The paper's summary: Nadia: So, the paper summarizes by showing these nine classes of failure modes and explains how they connect through cross-cutting patterns. They’re not just a list; it’s showing how one bad decision in scoping can lead to problems down the line in testing or deployment.
Elias: They identified four main patterns that repeat across those nine classes, like security being evaluated in the negative, and causes happening before the evaluation even starts because decisions are made way too early.
Priya: It really hammers home that there’s a huge gap between what you write down on paper and what actually happens when you deploy something in the real world. They point out that often, we’re just dealing with a "conformant-on-paper versus secure-as-deployed" divergence.
Nadia: And this framework they propose is meant to be actionable for product teams, giving them concrete steps—like making sure the device is drivable as a client and making failures inducible—to stop those divergences before evaluation even begins.
Elias: So, if you look at the summary, it’s not just a critique of how evaluations stall; it’s proposing a concrete design-for-evaluability framework that product teams can adopt to shorten the evaluation time and cost.
The paper's improvements: Nadia: Now for what they suggest as improvements, which is basically turning this diagnostic literature into something constructive. They focus on making these practices specific to those three moments where design and development practices attach.
Elias: They suggest things like making sure when you scope a Security Target, you draw the boundary from evidence instead of just ambition, and running a gap analysis before committing to the Protection Profile. That sounds like a practical way to avoid those initial errors.
Priya: I see them suggesting that product teams need to make sure their design provides things like making the device drivable as a client and making failures inducible, which is important because if you can’t make failures happen on purpose, you can’t really test for them.
Nadia: Plus, they push for better error handling: when a device rejects something, it should have to say why in a way that an evaluator or administrator can actually capture that information in a log line or error code. That fixes the issue of unclear refusals.
Elias: They also touch on making sure claims are backed up by real proof, like requiring an inventory of cryptography before claiming it and scheduling entropy assessments first to stop relying on inherited components without ownership.
Conclusion: Nadia: So, wrapping this up, the big implication of the SoK paper is that most of these failure modes are preventable cheaply and unilaterally by product teams themselves rather than waiting for expensive laboratory failures.
Elias: They argue that seven out of the nine classes trace directly to decisions a product team controls, and every practice mentioned can be adopted by one vendor alone, meaning it’s about existing engineering discipline reaching the artifacts.
Priya: From my side, it means we need to stop treating certification as something that happens *after* everything is built and just check a box; it needs to be woven into the entire design process proactively.
Nadia: It does, and the SoK paper gives us the structure to see exactly where those weave points are located across all those lifecycle stages.
Elias: Ultimately, this work points out that while we have these great frameworks, the model itself still has open problems when products move toward continuous delivery or cloud deployment because they strain that fixed configuration evaluation model.
Priya: I agree, and the paper is clear about its limitations: it doesn't solve the problem of what actually happens when certification is applied to a product that’s constantly moving, like in a continuous delivery pipeline.
Nadia: So we leave it there for this discussion on SoK: Failure Modes in Common Criteria Product Evaluation—A Taxonomy and Design-for-Evaluability Guidance. We’ll be back next week with another piece on how AI agents are actually being evaluated by security researchers.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails