System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "System-Level Optimization Beyond Cryptographic Kernels".
Nadia: Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetickernel improvements, assembly tuning, register allocation, and instruction scheduling.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at "System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7," which really suggests they're going beyond just making the math run faster inside the processor core. This paper looks at how we can squeeze performance out of the whole system, not just the arithmetic itself.
Elias: I agree; it's interesting because it moves away from just tuning assembly and focuses on how that code interacts with memory and hardware. They are investigating gains from things like memory hierarchy utilization, tightly coupled memory placement, peripheral integration, and deterministic public-data reuse one.
Priya: From a privacy research angle, I'm curious about what those system-level tweaks actually mean for data handling. If they are caching public data or moving it around on the chip, we need to understand how that state management works because that directly impacts any potential side channels.
Nadia: That’s a fair concern, Priya; the paper tackles repeated public-data derivation as a major remaining cost in optimized ML-KEM and designs wire-format-preserving techniques for ML-KEM-five hundred twelve ML-KEM-seven hundred sixty-eight and ML-KEM-one thousand twenty-four three. They found that choosing a specific public data reuse profile can cut encapsulation cycles by up to seventy-four point six percent or decapsulation cycles by up to fifty-eight point eight percent three.
Elias: That level of reduction is significant when you consider the computational load these lattice-based schemes carry, which is why this system-level optimization matters so much for real embedded deployment one. It really highlights that the performance bottleneck isn't just in the math itself.
The paper's summary: Nadia: So, to summarize what this paper is actually doing, they propose a two-stage optimization framework that complements the instruction-level tuning with a systematic evaluation of the execution system. They break it down into an "Optimized Kernel" stage and then an "Optimized Embedded System" stage two.
Elias: That framework is really neat because it formalizes how you should approach embedded optimization by first perfecting the math, and then optimizing the environment around that math, which includes choices about memory placement and hardware integration two. They tested specific system-level knobs like instruction tightly coupled memory placement, data tightly coupled memory placement for public constants, hardware true-random-number-generator integration, and clock configuration across a controlled sweep from twenty-four MHz to two hundred sixteen MHz two.
Priya: When you talk about those specific knobs, like placing frequently executed code into Instruction Tightly Coupled Memory or moving public constants to Data Tightly Coupled Memory, what kind of practical impact does that have on the actual data flow we see? Does it mean fewer stalls waiting for external memory access?
Nadia: It absolutely means reducing latency from those external memory accesses, Priya. They identified several classes of optimization opportunities, such as hot-code placement to reduce instruction fetch overhead and hot-data placement to speed up data access two. This moves the focus from just the algorithm's complexity to how the hardware executes it efficiently.
Elias: And they also looked at integrating hardware true-random-number generators into the process, which is a specific example of hardware offloading that bypasses software overhead two. It shows they aren't just talking abstract theory but testing concrete platform features on the Arm Cortex-M7 architecture.
The paper's improvements: Nadia: So to wrap up this paper, the main conclusion is that even after you've done all the low-level arithmetic optimization, there are still substantial gains possible by systematically optimizing the execution system—memory, hardware interactions, and clock settings two. They also developed practical guidelines for selecting deployment-oriented profiles based on these system-level optimizations.
Elias: The implication here is that for deploying post-quantum cryptography in resource-constrained environments, focusing only on the arithmetic kernel isn't enough; you have to treat the entire operational lifecycle as a target for optimization two. They pointed out that without modifying the algorithm or standardized wire formats, these system-level evaluations still showed gains, with some public data reuse profiles achieving reductions of up to seventy-four point six percent in cycles one.
Priya: I think what this paper really tells us is that performance isn't just about the raw mathematical complexity; it’s deeply tied to the physical constraints of where you put your memory and how frequently you can reuse public values during sequential operations three. It gives a clearer picture of how much overhead we can realistically shave off in real hardware.
Nadia: Agreed, Priya. This work on "System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7" really highlights that for embedded systems, the physical execution environment is just as important as the logic running inside it two. Elias, what do you think about the future work they suggested?
Elias: They suggest that their methodology transfers as a set of engineering questions where measured rankings depend heavily on the specific platform and workload, which means future work needs to focus on creating more generalized selection criteria for these system-level optimizations two.
Priya: I wonder if there's a path to measuring the privacy impact of these different optimization profiles, since they are all tied back to how public data is handled across the lifecycle? That seems like a logical next step for any researcher looking at this.
Conclusion: Nadia: So we've looked at how optimizing things beyond just the core arithmetic really opens up new avenues for performance in embedded systems, especially when looking at this paper, "System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7."
Elias: Exactly. It shows that even after you nail the instruction scheduling and register allocation, there's still a whole layer of execution environment choices—memory placement, clock settings—that can make a measurable difference in wall-clock time.
Priya: From my side, I'm really focused on what the data actually shows regarding public data reuse; seeing those cycle reductions suggest that caching deterministic values is a very practical way to manage performance without necessarily introducing major new security vectors, provided the state management is handled correctly three.
Nadia: That's a solid point about practical application, Priya. And Elias was right before; the gains are substantial, especially when you look at how much faster encapsulation and decapsulation become with those specific reuse profiles. It really demonstrates that system-level tuning isn't just academic theory anymore for these types of cryptographic workloads.
Elias: I agree; the fact that they tested clock configuration experiments shows that even tweaking the processor frequency can reduce wall-clock latency when cycle counts stay stable, which is a key operational insight for anyone designing low-power edge devices two.
Priya: It's fascinating how these specific hardware choices, like endpoint-selected ITCM placement for the Keccak permutation and NTT kernels, translate into tangible speedups that depend entirely on how well the system is mapped to the underlying hardware structure two.
Nadia: And I think that's what makes this work so exciting; it gives us a clearer roadmap for optimizing future AI inference accelerators or any resource-constrained device where performance hinges on balancing computation with memory access latency.
Elias: It certainly provides a strong foundation for future work, though the paper does flag that these optimization choices are highly dependent on the specific deployment profile you choose, meaning there isn't one universal setting two.
Priya: That dependency on the deployment profile is something I think we need to keep watching closely because it tells us exactly where the trade-offs between speed and memory usage live for real-world applications.
Nadia: Well, that's our time on this paper. It’s clear that optimizing the entire execution system, not just the arithmetic kernel, is a vital step for making embedded AI systems truly efficient two.
Mahmoud Abdelhafeez Sayed, Mostafa Taha, Gurp Nijjer
Department of Systems and Computer Engineering, Carleton University · Quantegra Technologies Inc.
cs.CR
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: 22 pages and 4 figures. Extended version to the paper in the Proceedings of FPS-2026: 19th International Symposium on Foundations & Practice of Security
Code: https://github.com/Carleton-SCI/ML-KEM-M7-SYS-OPT
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetickernel improvements, assembly tuning, register allocation, and
Key concepts
- Two-Stage Optimization Framework
- A two-step approach to optimization: first, tuning the core arithmetic instructions (Instruction-Level); second, optimizing how that code runs in the system environment (System-Level). This ensures both fast computation and efficient execution within hardware constraints.
- Hot-code Placement
- Placing frequently executed code regions into low-latency instruction memories. This reduces the time spent fetching instructions from slower external memory, which speeds up the overall process by minimizing instruction fetch overhead.
- Public-Data Reuse
- Caching deterministic values derived only from public information to avoid recalculating them repeatedly. This trades extra on-chip storage for massive reductions in computation cycles during key operations like encapsulation and decapsulation.
Terminology
Summary
Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetickernel improvements, assembly tuning, register allocation, and instruction scheduling. Using the Module-Lattice-Based Key-Encapsulation Mechanism (ML-KEM) on an Arm Cortex-M7 as a case study, this paper examines the additional gains available from memory-hierarchy utilization, tightly coupled memory placement, peripheral integration, clock configuration, and deterministic public-data reuse.
The gist: Substantial deployment gains remain after arithmetic-kernel optimization; evaluated profiles without auxiliary public state reduce cycles by up to 2.5%, while a selected public-data-reuse profile reduces encapsulation and decapsulation cycles by up to 74.6% and 58.8%, respectively.
The Two-Stage Optimization Framework
The paper proposes a two-stage methodology for embedded optimization that complements instruction-level tuning with systematic evaluation of the execution system, including memory placement, hardware integration, operating points, and deterministic public state across operation and key lifecycles. This framework organizes optimization into two stages: Instruction-Level Optimization and System-Level Optimization. The first stage focuses on reducing local computational cost through techniques such as arithmetic design, transform restructuring, instruction scheduling, register allocation, software pipelining,
resulting in an Optimized Kernel.
The second stage addresses the execution and reuse of those kernels within a deployment through choices like memory placement, peripheral integration, operating-point selection, and management of deterministic state across API calls,
leading to an Optimized Embedded System.
System-Level Optimization Classes
The optimization opportunities encountered in embedded cryptographic workloads are grouped into several recurring system-level optimization classes. These include:
-
Hot-code placement:
places frequently executed code regions into lowlatency instruction memories to reduce instruction-fetch overhead.
-
Hot-data placement:
relocates frequently accessed data structures into fast data memories to reduce access latency and improve execution predictability.
-
Public-data reuse:
trades additional storage for reduced recomputation by caching deterministic values derived exclusively from public information.
-
Hardware offloading:
utilizes platform peripherals or accelerators to reduce software overhead and free processor resources.
-
Clock and operating-point optimization:
adjusts processor and memory operating parameters to reduce wall-clock execution time or improve energy efficiency.
Specific System-Level Optimization Directions
The paper evaluates several specific system-level optimization directions on the ML-KEM implementation on the Arm Cortex-M7 baseline. These directions include:
: Endpoint-selected ITCM placement:
This corresponds to Hot-code placement
and involves relocating frequently executed code regions into lowlatency instruction memory. The evaluation focuses on Endpoint-selected ITCM,
which is represented by ITCM Placement of Hot Code,
where the Cortex-M7's 16 KiB ITCM RAM is used for groups like the Keccak permutation, NTT, and matrix-accumulation kernels.
: DTCM NTT constants:
This direction addresses Hot-data placement
by copying public twiddle-factor tables from Flash into DTCM at startup to remove loads from the Flash path. The cost is noted as startup copying and additional random-access memory (RAM).
: Hardware TRNG Integration:
This evaluates the Hardware offloading
class, using the STM32 peripheral as a source for fresh randomness instead of a deterministic software fallback.
: Clock and Operating-Point Optimization:
This involves clock configuration experiments,
where changing the processor clock frequency is treated as an operating-point optimization to reduce wall-clock execution time, noting that wall-clock latency decreases approximately with increasing processor-clock frequency (HCLK) when cycle counts remain stable.
Public-Data Reuse Techniques
The paper identifies repeated public-data derivation as a major remaining cost and designs wire-format preserving techniques for ML-KEM. The optimization class is represented by:
-
Cached A: Captures a
serialized, multiplication-ready representation
of the public matrix A, which is expected to yieldlargest isolated gains
in encapsulation and decapsulation cycles but requires auxiliary public state. -
Cached H(pk): Caches the hash of the public key, which reduces encapsulation cycles by 12.49–16.35% for only 32 B of state, while requiring
+32
bytes of auxiliary state. -
Parsed Public-Key Cache: Stores the decoded public representation to avoid repeated parsing and preparation during encapsulation and decapsulation, showing a smaller effect but requiring
2–4 KiB.
Deployment-Oriented Profile Selection
The final stage involves deriving practical system-level optimization classes, selection criteria, and deployment-oriented profile-selection guidelines.
The study identifies several balanced profiles.
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7.
While the paper focuses on post-quantum cryptography (PQC) implementation optimization rather than general AI system architecture, the principles of its two-stage optimization framework are highly transferable to optimizing resource-constrained AI/ML systems.
Here are specific improvements and what those improved AI systems could achieve:
I. Improvements Based on the Two-Stage Optimization Framework
The paper's core contribution is moving optimization beyond instruction-level kernels to the execution system (memory, hardware, clock). Applying this methodology to an embedded AI accelerator or edge device would yield significant gains.
-
Improvement: Implement a two-stage optimization pipeline: first optimize the core neural network arithmetic (the
cryptographic kernel
equivalent), and second, systematically evaluate the surrounding execution environment (memory placement, peripheral integration, clock configuration). -
Improvement: Develop systematic classification and selection criteria for system-level optimizations tailored to AI workloads. This includes defining classes like
Hot-code Placement
(placing frequently run inference layers in fast local memory),Hot-data Placement
(relocating weights and activations to low-latency on-chip memory), andPublic/Activation Reuse.
-
Improvement: Design techniques for deterministic reuse of intermediate results or public parameters across sequential inference steps, analogous to the ML-KEM's public data reuse.
II. Specific System-Level Optimization Techniques (Mapping PQC Concepts to AI)
Based on Section 3 and Section 4, here are specific optimizations applicable to an AI system:
-
Improvement: Implement selective placement of frequently executed inference layers (e.g., the first few convolution layers or attention mechanisms) into low-latency Instruction/Data Memory (ITCM/DTCM equivalents).
-
Improvement: Employ
Hot-Data Placement
for model weights and intermediate feature maps, moving them from slower external memory (Flash/DRAM) to fast on-chip SRAM to reduce access latency and improve execution predictability. -
Improvement: Integrate hardware offloading for specific compute-intensive operations (e.g., matrix multiplication or polynomial transforms in the PQC analogy) using dedicated accelerators or specialized DSP features, reducing software overhead.
-
Improvement: Systematically evaluate clock configuration (Operating-Point Selection) to determine the optimal processor frequency that balances execution speed against power/thermal constraints for a given workload mix (e.g., high-frequency for low latency on edge devices, lower frequency for battery life).
-
Improvement: Implement caching strategies for frequently accessed model parameters or activation outputs (Public-Data Reuse), trading additional on-chip memory (RAM) and key generation/preparation overhead for reduced repeated computation during sequential inference.
III. Enhanced AI System Capabilities
By applying these improvements, the resulting AI systems can achieve the following specific capabilities:
-
Reduced Inference Latency: By optimizing instruction and data movement to fast local memory (ITCM/DTCM), the system can drastically reduce the wall-clock time required for a single inference pass, achieving sub-millisecond latency even on resource-constrained edge devices (as demonstrated by the PQC results).
-
Increased Throughput on Fixed Hardware: Effective memory placement and hardware offloading allow the system to sustain a higher number of inferences per second by minimizing memory access bottlenecks, effectively increasing the utilization of the limited on-chip resources.
-
Improved Energy Efficiency: By using clock scaling and selecting appropriate operating points based on workload characteristics, the system can achieve better performance-per-watt than simply running at maximum frequency across all operations.
-
Efficient Handling of Sequential Tasks: The ability to cache public/derived states allows the AI system to perform complex, multi-step tasks (like chained sequential inference or dynamic routing) much faster by avoiding redundant computations of intermediate results.
In summary, the improved AI system will be a highly optimized embedded accelerator capable of delivering low-latency, high-throughput inference on resource-constrained hardware by treating the entire execution environment—not just the arithmetic—as an optimization target.
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs