SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance

arXiv:2610.10644 · cs.CR, cs.SE · Submitted 2026-10-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance".

Elias: The gist The paper presents a systematization of knowledge (SoK) regarding failure modes in Common Criteria product evaluation,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So, looking at this paper again, "SoK: Failure Modes in Common Criteria Product Evaluation—A Taxonomy and Design-for-Evaluability Guidance," the point is that certification isn't just some checkbox thing anymore; it’s a process full of hidden pitfalls.

Elias: The authors are taking that gap between what a product claims it is and what the evaluation actually finds, and they're trying to close it by creating this structured taxonomy of failure modes. It’s about making the evaluation process itself more predictable.

Priya: What I find interesting is how they frame this entire thing around a design-for-evaluability framework, which means shifting the focus from just fixing problems after testing to preventing those problems in the very early stages of product development.

Nadia: Right, so instead of waiting for a failure to happen during evaluation, the paper suggests we build practices into the design phase—things like making sure you’re drawing your security boundary based on actual evidence rather than just what you *hope* it's going to be.

Elias: That ties into how they structure it: they organize failures across nine lifecycle classes, which follow the path an evaluation actually takes, from scoping the Security Target right up to what happens after you get that certificate.

The paper's summary: Nadia: So, the paper summarizes by showing these nine classes of failure modes and explains how they connect through cross-cutting patterns. They’re not just a list; it’s showing how one bad decision in scoping can lead to problems down the line in testing or deployment.

Elias: They identified four main patterns that repeat across those nine classes, like security being evaluated in the negative, and causes happening before the evaluation even starts because decisions are made way too early.

Priya: It really hammers home that there’s a huge gap between what you write down on paper and what actually happens when you deploy something in the real world. They point out that often, we’re just dealing with a "conformant-on-paper versus secure-as-deployed" divergence.

Nadia: And this framework they propose is meant to be actionable for product teams, giving them concrete steps—like making sure the device is drivable as a client and making failures inducible—to stop those divergences before evaluation even begins.

Elias: So, if you look at the summary, it’s not just a critique of how evaluations stall; it’s proposing a concrete design-for-evaluability framework that product teams can adopt to shorten the evaluation time and cost.

The paper's improvements: Nadia: Now for what they suggest as improvements, which is basically turning this diagnostic literature into something constructive. They focus on making these practices specific to those three moments where design and development practices attach.

Elias: They suggest things like making sure when you scope a Security Target, you draw the boundary from evidence instead of just ambition, and running a gap analysis before committing to the Protection Profile. That sounds like a practical way to avoid those initial errors.

Priya: I see them suggesting that product teams need to make sure their design provides things like making the device drivable as a client and making failures inducible, which is important because if you can’t make failures happen on purpose, you can’t really test for them.

Nadia: Plus, they push for better error handling: when a device rejects something, it should have to say why in a way that an evaluator or administrator can actually capture that information in a log line or error code. That fixes the issue of unclear refusals.

Elias: They also touch on making sure claims are backed up by real proof, like requiring an inventory of cryptography before claiming it and scheduling entropy assessments first to stop relying on inherited components without ownership.

Conclusion: Nadia: So, wrapping this up, the big implication of the SoK paper is that most of these failure modes are preventable cheaply and unilaterally by product teams themselves rather than waiting for expensive laboratory failures.

Elias: They argue that seven out of the nine classes trace directly to decisions a product team controls, and every practice mentioned can be adopted by one vendor alone, meaning it’s about existing engineering discipline reaching the artifacts.

Priya: From my side, it means we need to stop treating certification as something that happens *after* everything is built and just check a box; it needs to be woven into the entire design process proactively.

Nadia: It does, and the SoK paper gives us the structure to see exactly where those weave points are located across all those lifecycle stages.

Elias: Ultimately, this work points out that while we have these great frameworks, the model itself still has open problems when products move toward continuous delivery or cloud deployment because they strain that fixed configuration evaluation model.

Priya: I agree, and the paper is clear about its limitations: it doesn't solve the problem of what actually happens when certification is applied to a product that’s constantly moving, like in a continuous delivery pipeline.

Nadia: So we leave it there for this discussion on SoK: Failure Modes in Common Criteria Product Evaluation—A Taxonomy and Design-for-Evaluability Guidance. We’ll be back next week with another piece on how AI agents are actually being evaluated by security researchers.

Punit Suketu Patel

cs.CR, cs.SE

Submitted: 2026-10-07

Updated: 2026-10-07

Comments: 25 pages, 1 figure, 1 table. Accepted at the Security Standardisation Research (SSR) Conference 2026; to appear in Springer LNCS. Author's submitted version, prior to peer-review revisions

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

The gist: The gist The paper presents a systematization of knowledge (SoK) regarding failure modes in Common Criteria product evaluation, drawing on recurring cross-vendor failures across nine lifecycle

Key concepts

Lifecycle Taxonomy
A structured classification system organizing recurring failure modes into nine distinct classes that map the entire product lifecycle, from initial security target scoping through post-certification configuration drift. Failures are only included if they recur across vendors and can be made actionable.
Design-for-Evaluability Framework
A set of practices product teams should adopt before evaluation starts. This framework focuses on ensuring the device is drivable as a client and that failures are easily inducible, such as providing clear error logs when a device rejects input.
Evidence Anchors by Failure Class
A detailed mapping showing which specific work units (like ASE or ATE) anchor particular failure modes. This links abstract problems to concrete engineering practices and identifies exactly where in the evaluation process a failure occurs, such as boundary redrawing or documentation errors.

Terminology

Summary

The gist The paper presents a systematization of knowledge (SoK) regarding failure modes in Common Criteria product evaluation, drawing on recurring cross-vendor failures across nine lifecycle classes to derive a design-for-evaluability framework for product teams.

Taxonomy and Root Cause Analysis

The paper organizes recurring failures into a lifecycle taxonomy spanning Security Target scoping, Protection Profile conformance, assurance evidence, cryptographic requirements, functional testing, vulnerability analysis, operational guidance, and post-certification configuration drift <ref:2610.10644#pg8> The taxonomy was not assembled from a blank page <ref:2610.10644#pg6> Its nine classes follow the path an evaluation actually walks, from scoping the Security Target to what happens after the certificate is issued <ref:2610.10644#pg4> Each class is anchored to the assurance family or CEM work units where that failure surfaces <ref:2610.10644#pg6> A failure mode earned a place in the taxonomy if it met three tests: it recurs across vendors and product categories rather than being one product’s bad day <ref:2610.10644#pg6> it maps to an identifiable work unit, so an evaluator can point to where it bites <ref:2610.10644#pg6> and it is actionable, meaning a product team could have prevented it before evaluation began <ref:2610.10644#pg6>.

Design for Evaluability Framework

The paper derives a design-forevaluability framework that product teams can apply before evaluation begins <ref:2610.10644#pg4> This framework is organized into three moments where practices attach <ref:2610.10644#pg6>. When the claim is scoped, teams must draw the boundary from evidence, not ambition <ref:2610.10644#pg6> and run a gap analysis before committing to the PP <ref:2610.10644#pg6>. What the design must provide includes making the device drivable as a client and making failures inducible <ref:2610.10644#pg6>. Furthermore, teams should make refusals speak when a device rejects something, have it say why, in a log line or error code that an evaluator or an administrator can capture <ref:2610.10644#pg6>.

Open Problems and Model Strain

The paper identifies open problems where the model itself is the cause <ref:2610.10644#pg6> These problems live past the prevention boundary, such as certifying software that moves due to continuous delivery or cloud deployment <ref:2610.10644#pg6>. The open question is what the unit of certification becomes when the product is a stream rather than an artifact <ref:2610.10644#pg6>. Additionally, there are questions regarding requirements for components that cannot be enumerated, such as Machine Learning components <ref:2610.10644#pg6>.

Cross-Cutting Patterns

The taxonomy reveals four patterns that repeat across the nine classes <ref:2610.10644#pg6>. The described product and the shipped one show a shape of a description on one side, a behavior on the other, and the failure living in the distance between <ref:2610.10644#pg6>. Security is evaluated in the negative, where testing involves staging forbidden things and watching them be refused for the right reason <ref:2610.10644#pg6>. The causes precede the evaluation, as decisions were made before the laboratory was engaged <ref:2610.10644#pg6>. Finally, a certificate cannot say what it knows, compressing all information into one bit plus documents that do not travel <ref:2610.10644#pg6>.

Conclusion

Most of this is preventable, cheaply and unilaterally <ref:2610.10644#pg6> Seven of the nine classes trace to decisions a product team controls, the laboratory’s playbook is public, and every practice in Section 6 can be adopted by one vendor alone <ref:2610.10644#pg6>. Design for evaluability asks for no new discipline, only for existing engineering discipline to reach the artifacts an evaluation consumes >. The model itself remains the subject of open problems <ref:2610.10644#pg6>.

A Evidence Anchors by Failure Class

Class Anchor Work-unit locus What it evidences

E1 ST 383-7-138, v2.1, revision history [4] ASE Boundary redrawn mid-evaluation: “Expanded the TOE boundary to address certifier ORs” E1 JBoss EAP 4.3.0 ST, BSI-DSZ-CC-0531 [33] ASE Revisions log the boundary moving in both directions: requirements removed, TOE components added E2 TD0592 [31] ASE CCL PP-mandatory local audit storage absent from a shipping product; resolved by scheme decision E2 TD0009 [26] ASE CCL Claimed cipher-suite selections in conflict with what implementations honor (WLAN AS PP) E3 FortiMail ST v1.13; ST VID 11253 [13, 25] ATE, prescribed activities [23] Amended STs on the record: testing halts on the affected requirements, the claim is amended, review re-opens E4 Practitioner FAQ [16] CEM document checks Laboratory–vendor iteration until each document carries the CEM’s checks, documented as the working norm E4 ICCC practitioner reports [2, 9] ADV, ALC inputs Documents written for the wrong job, reported from the evaluation floor ADV/ALC work units; ATE IND evidence; ETR assembly Documents treated as a byproduct of the work, not a product of it Write documents from the build; plan evidence capture before testing E5 NIAP Policy Letter No. 5 [24] FCS; CAVP gate Every claimed algorithm with a NIST validation program must hold a CAVP certificate; scope mismatches are hard stops E5 SP 800-90B [36] Entropy assessment Raw-data demands that exceed what shipping devices store or expose E5 TD0057 → TD0130 [27, 28] FCS CKM key destruction Public iteration arc on one requirement: zeroization true at the API, false on wear-leveled media E6 NDcPP Supporting Document [23] ATE IND activities The prescribed per-requirement test activities whose nominal completeness E6 measures against E6 TD0190 [29] ATE IND Failure-state requirements that cannot be induced without vendor test builds E7 CEM [7] AVA VAN The actual mandate: public-source survey plus penetration testing at basic attack potential—narrower than buyers read E7 ROCA [22] bound Inexpensive attacks on certified payment devices, demonstrated post-certification E7, E9 Certificate corpus [15] Postcertification record Disclosed vulnerabilities mapped onto certified products at scale E8 CC documentation services [34, 2] AGD Guidance authored away from the hardware by writers and consultants—industry practice on the record since 2001 AGD PRE / AGD OPE; bring-up Wrong steps, lost on reboot, or unread supplement Written away from the hardware; never tested like code Fresh-device walkthrough; explicit persistence; posture status command A supported status command that attests the certified posture would do more than any amount of better prose E9 Postcertification drift Product and deployment move; the certificate stands still After issuance; assurance continuity Maintenance Version-anchored maintenance: the mechanism E9 measures against exploit-speed patching

--- Page 24 ---

A Evidence Anchors by Failure Class

Class Anchor Work-unit locus What it evidences

E1 ST 383-7-138, v2.1, revision history [4] ASE Boundary redrawn mid-evaluation: “Expanded the TOE boundary to address certifier ORs” E1 JBoss EAP 4.3.0 ST, BSI-DSZ-CC-0531 [33] ASE Revisions log the boundary moving in both directions: requirements removed, TOE components added E2 TD0592 [31] ASE CCL PP-mandatory local audit storage absent from a shipping product; resolved by scheme decision E2 TD0009 [26] ASE CCL Claimed cipher-suite selections in conflict with what implementations honor (WLAN AS PP) E3 FortiMail ST v1.13; ST VID 11253 [13, 25] ATE, prescribed activities [23] Amended STs on the record: testing halts on the affected requirements, the claim is amended, review re-opens E4 Practitioner FAQ [16] CEM document checks Laboratory–vendor iteration until each document carries the CEM’s checks, documented as the working norm E4 ICCC practitioner reports [2, 9] ADV, ALC inputs Documents written for the wrong job, reported from the evaluation floor ADV/ALC work units; ATE IND evidence; ETR assembly Documents treated as a byproduct of the work, not a product of it Write documents from the build; plan evidence capture before testing E5 NIAP Policy Letter No.

Improvements for AI systems

  1. Improve Security Target Scoping by implementing a mandatory evidence production capability check before contract finalization, directly addressing E1 failures where a boundary drawn too broad commits the vendor to producing design evidence and test access for components it does not control.

  2. Enhance Protection Profile Conformance checks by requiring a pre-evaluation requirement-by-requirement gap analysis against the exact PP version to prevent claims that are conformant on paper, different in practice, as noted in E3.

  3. Strengthen Cryptographic Assurance by mandating an inventory the cryptography before claiming it and scheduling the entropy assessment first, ensuring that claims do not rely on inherited components where inheritance without ownership is identified as a root cause for E5 failures.

  4. Implement Drivable-as-Client interfaces to enable clients to initiate protocol connections with test-controlled parameters, specifically addressing E6 by ensuring client-role requirements cost what server-role ones do.

  5. Develop mechanisms for Inducible Failure Conditions to ensure that testing exercises the shipping build and not a modified one, thereby preventing failures where the test leans on a vendor-supplied test build or debug hook (E6).

  6. Integrate Posture Status Commands into the device's operation to make E9 drift visible, allowing administrators to check am I in the evaluated configuration right now? which repairs E8’s unread supplement.

Related papers