Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges
summary
The gist
The gist: This work presents a solution that allows for the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone.
In short
The work presents a method to run unaltered AI models inside edge Trusted Execution Environments (TEEs) like OP-TEE for Arm TrustZone. It achieves this by compiling AI models to WebAssembly (Wasm) and using the WebAssembly Micro Runtime (WAMR) within OP-TEE. This allows secure execution of encrypted models, protecting intellectual property while addressing TEE constraints.
Key concepts
- Arm TrustZone
- This is a hardware security extension that divides a device's processor into two worlds: the Normal World and the Secure World. The Secure World has special access to sensitive resources, ensuring that critical operations and data remain protected from potentially compromised applications running in the Normal World.
- WebAssembly (Wasm)
- Wasm is a portable compilation target for existing applications. It allows developers to compile AI models into this format, which can then be executed securely inside the TEE. This portability solves the problem of having to heavily modify AI software just to run it in a constrained TEE environment.
- WAMR (WebAssembly Micro Runtime)
- WAMR is the specific runtime used within OP-TEE to execute WebAssembly code. It acts as a translator, taking Wasm instructions and converting them into native operations that the OP-TEE system can understand and run securely on the hardware.
Terminology used across episodes
This episode discusses
- Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges · Paper Radio
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
The paper
Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges · Read on arXiv
Friedrich Vandenberghe, Lachlan Gunn, Bruno Volckaert, Merlijn Sebrechts
IDLab, Department of Information Technology Ghent University - imec, Belgium · Aalto University, Finland
AI models on edge hardware contain important intellectual property (IP), but an adversary can steal it when they achieve root access. Trusted Execution Environments (TEE) like Arm TrustZone protect against these Operating System (OS) level attacks. However, they are challenging to use. More precisely, it is difficult to run unaltered applications inside a TEE. This work presents a solution that allows the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone. Additionally, this work provides an AI model distributor that encrypts the WebAssembly binary and places the encryption key in one of the device's fuses. This way, only the WAMR Trusted Application (TA) in OP-TEE can decrypt and execute the AI model. A thorough evaluation of our solution shows that it incurs an additional overhead of 22% in comparison to an application manually ported to OP-TEE, while also facilitating the execution of unaltered AI models with an additional inference latency of 6% without significant porting effort. Overall, there is a pressing need to safeguard the IP of AI models and this work shows that there is a real promise in WebAssembly, but there remain some practical challenges.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Protecting CPU AI On Edge TEEs".
Elias: The gist: This work presents a solution that allows for the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at this paper, "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges." It’s talking about how to actually run existing AI models on edge hardware without giving the intellectual property away when someone gets root access.
Elias: Yeah, it tackles the problem that TEEs like Arm TrustZone are great for security but they make running unaltered applications inside them pretty hard without a ton of rewriting stuff. This paper proposes using WebAssembly to make that possible in OP-TEE.
Priya: What I wonder is how much of this protection actually translates into real confidentiality when you're dealing with these kinds of edge devices where things are always moving around.
Nadia: Exactly, Priya, and the paper suggests a specific way to do it: they compile those AI models into WebAssembly binaries first. Then, they encrypt those binaries and lock the key away in a physical fuse on the device itself so only the Trusted Application can unlock and run them.
Elias: And that whole setup relies on using OP-TEE with its WebAssembly Micro Runtime, or WAMR, to handle translating those WebAssembly instructions into what the TEE can actually execute natively.
Priya: So, if we’re talking about the real world, this means they're taking models that are already built and just wrapping them in this secure format instead of having to rebuild the entire application for the TEE.
Nadia: That's right. The paper is focused on answering a few specific research questions, like how WebAssembly can recompile existing AI models for Arm TrustZone without needing any changes to the original model itself.
Elias: And they also look at performance overhead, which is something everyone cares about when you're dealing with edge devices where resources are tight. They specifically ask what the performance overhead is when running these WebAssembly applications inside Arm TrustZone compared to native code.
Priya: I mean, for someone just driving or cooking, how does that translate? Is it a noticeable slowdown in startup time or during actual use?
Nadia: Well, the evaluation shows they found an additional overhead of twenty-two percent when comparing their solution to an application that was manually ported directly into OP-TEE.
Elias: That’s a significant number, so you can see it's not free. But they also noted that for inference latency, specifically for a MobileNetV2 model, they saw a mean slowdown of twenty-one point seven microseconds, which is about five point six one percent slower than running it in the WAMR environment inside OP-TEE.
Priya: So the numbers show it adds some friction to the operation compared to a perfectly ported application. What about the security side? How robust is this encryption mechanism they’re suggesting?
Nadia: The security model is pretty strict; model confidentiality is maintained because you can’t get the decryption key, which they state only exists within the Secure World and can be read by the WAMR Trusted Application.
Title and authors: Elias: They also talk about platform confidentiality, ensuring that the OP-TEE OS verifies any Trusted Application being loaded against a public key bundled with it to stop anyone from loading a custom application in there.
Priya: That addresses the risk of someone loading their own malicious code into the TEE, which is a big win for protecting that local AI model IP.
Nadia: But then they also flagged some clear limitations. They mentioned there are restrictions because of the limited support for current workloads and specifically a limited set of supported ONNX operators, which means you can't just run any random model format.
Elias: And another practical challenge they point out is that the solution currently lacks GPU acceleration, which definitely limits how fast these AI models can actually perform their tasks on the hardware.
Priya: So what does this mean for someone listening? It means you get a secure way to keep your AI model safe from theft when it’s running on your device, but you have to be very careful about which specific models you use and how much speed you expect.
Nadia: That's the summary of the trade-off: security and portability in exchange for some overhead and feature limitations right now. We need to see more work on expanding that AI model support, maybe by switching to an ONNX runtime, because that’s where they feel the biggest opportunity is left.
Elias: If we look ahead, they suggest future work will involve advancing the associated WASI neural network proposal and switching to a proper ONNX runtime. That would hopefully open up those performance gaps they are currently seeing in inference speed.
Priya: I’m just hoping that as the ecosystem matures, these kinds of solutions become more practical for everyday edge computing needs rather than just high-end research prototypes.
Nadia: That’s the direction we need to watch. So, to wrap up on this paper, "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges," it shows a way to execute unaltered AI models in OP-TEE for Arm TrustZone using encrypted WebAssembly binaries.
Elias: It’s a feasible approach that addresses the difficulty of running unaltered applications inside TEEs by leveraging WAMR and a custom distribution method where the encryption key is stored in device fuses.
Priya: And while it adds a twenty-two percent overhead compared to native ports, it still offers model confidentiality against root access attacks when dealing with ONNX formatted models.
Nadia: The real constraint is that the current implementation is restricted by the limited set of supported ONNX operators and the lack of GPU acceleration, so future work needs to focus on broadening those capabilities.
Elias: That’s where they are heading, aiming for an ONNX runtime and advancing that WASI neural network proposal to fix those performance issues.
Priya: So, for anyone listening who cares about the practical application, this paper gives us a concrete roadmap of what needs to be done next to make this technology truly ready for widespread edge deployment.
The paper's summary: Nadia: So, we’re looking at this paper that shows how you can run existing AI models on edge hardware without giving away the intellectual property when someone gets root access.
Elias: It tackles the problem that TEEs like Arm TrustZone are great for security but they make running unaltered applications inside them pretty hard without a ton of rewriting stuff.
Priya: What I wonder is how much of this protection actually translates into real confidentiality when you're dealing with these kinds of edge devices where things are always moving around.
Nadia: Exactly, Priya, and the paper suggests a specific way to do it: they compile those AI models into WebAssembly binaries first. Then, they encrypt those binaries and lock the key away in a physical fuse on the device itself so only the Trusted Application can unlock and run them.
Elias: And that whole setup relies on using OP-TEE with its WebAssembly Micro Runtime, or WAMR, to handle translating those WebAssembly instructions into what the TEE can actually execute natively.
Priya: So, if we’re talking about the real world, this means they're taking models that are already built and just wrapping them in this secure format instead of having to rebuild the entire application for the TEE.
Nadia: That's right. The paper is focused on answering a few specific research questions, like how WebAssembly can recompile existing AI models for Arm TrustZone without needing any changes to the original model itself.
Elias: And they also look at performance overhead, which is something everyone cares about when you're dealing with edge devices where resources are tight. They specifically ask what the performance overhead is when running these WebAssembly applications inside Arm TrustZone compared to native code.
Priya: I mean, for someone just driving or cooking, how does that translate? Is it a noticeable slowdown in startup time or during actual use?
Nadia: Well, the evaluation shows they found an additional overhead of twenty-two percent when comparing their solution to an application that was manually ported directly into OP-TEE.
Elias: That’s a significant number, so you can see it's not free. But they also noted that for inference latency, specifically for a MobileNetV2 model, they saw a mean slowdown of twenty-one point seven microseconds, which is about five point six one percent slower than running it in the WAMR environment inside OP-TEE.
Priya: So the numbers show it adds some friction to the operation compared to a perfectly ported application. What about the security side? How robust is this encryption mechanism they’re suggesting?
Nadia: The security model is pretty strict; model confidentiality is maintained because you can’t get the decryption key, which they state only exists within the Secure World and can be read by the WAMR Trusted Application.
Elias: They also talk about platform confidentiality, ensuring that the OP-TEE OS verifies any Trusted Application being loaded against a public key bundled with it to stop anyone from loading a custom application in there.
Priya: That addresses the risk of someone loading their own malicious code into the TEE, which is a big win for protecting that local AI model IP.
Nadia: But then they also flagged some clear limitations. They mentioned there are restrictions because of the limited support for current workloads and specifically a limited set of supported ONNX operators, which means you can't just run any random model format.
Elias: And another practical challenge they point out is that the solution currently lacks GPU acceleration, which definitely limits how fast these AI models can actually perform their tasks on the hardware.
Priya: So what does this mean for someone listening? It means you get a secure way to keep your AI model safe from theft when it’s running on your device, but you have to be very careful about which specific models you use and how much speed you expect.
Nadia: That's the summary of the trade-off: security and portability in exchange for some overhead and feature limitations right now. We need to see more work on expanding that AI model support, maybe by switching to an ONNX runtime, because that’s where they feel the biggest opportunity is left.
The paper's improvements: Nadia: So, we’re looking at what the authors suggest as improvements for this AI model execution setup on edge TEEs using WebAssembly.
Elias: They are basically saying they need to move past the current limitations because right now it’s too restrictive for real-world AI.
Priya: What exactly are they suggesting changes that would make this more useful for people actually building these systems?
Nadia: They're focusing on broadening the support beyond just a limited set of ONNX operators, which means they need to get better at handling different types of math operations in the AI models.
Elias: Yeah, and they point out that since it’s currently missing GPU acceleration, that’s a big thing because it limits how fast these models can actually perform their tasks on the hardware.
Priya: So if you want to run a faster model, this setup isn't going to give you that speed right now because of the hardware side?
Nadia: Exactly. The paper suggests future work needs to involve switching to a proper ONNX runtime, which should help with that performance issue.
Elias: And they also mention advancing the associated WASI neural network proposal, which is supposed to open up more ways for AI workloads to run securely inside the TEE environment.
Priya: So what does this mean for someone who just listens to the show? It means that for now, you can only use a narrow range of models, and you’re looking at some speed bumps.
Nadia: That’s the reality. They are calling out that the solution is currently restricted by memory management functions not part of the GlobalPlatform API, which they say OP-TEE Core provides as extensions instead.
Elias: So even if you get better operator support and a runtime, you still have to deal with those extra layers of complexity from the TEE core itself.
Priya: That sounds like it puts a lot of work on the developers who are trying to make these things usable for everyone, not just high-end research prototypes.
Nadia: Right. The authors are basically laying out a roadmap where they need to focus on that ONNX runtime and the WASI proposal if they want this thing to actually scale up beyond what it is today.
Conclusion: Tom: So we’re wrapping up on this paper, "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges," which shows how you can run existing AI models on edge hardware without giving away the intellectual property when someone gets root access.
Nadia: It basically confirms that WebAssembly combined with OP-TEE is a way to get unaltered AI models running inside secure environments like Arm TrustZone.
Elias: The main point here is using that custom distribution method where the encryption key stays in device fuses so only the trusted application can unlock and execute the model.
Priya: So what does this actually change for someone who just listens to the show? It means they can protect their local AI models from being stolen when they're running on their phone or device without needing to completely rewrite everything.
Nadia: Right. But we have to remember those performance trade-offs, like that twenty-two percent overhead compared to a manually ported application.
Elias: And the inference latency still shows a slowdown for models like MobileNetV2, which is about five point six one percent slower in the TEE WAMR environment.
Priya: I just want to focus on what the data actually shows: it’s secure against root access, but it’s not perfectly fast either right now.
Nadia: That's the constraint they are flagging—the limited support for specific ONNX operators and the lack of GPU acceleration are big roadblocks for what this can do today.
Elias: They are setting a clear path forward by saying future work has to involve switching to an ONNX runtime and advancing that WASI neural network proposal.
Priya: So, the implication is that this technology is feasible for protecting AI IP, but we need more development on the toolset to make it practical for real-world performance.
Nadia: Exactly. It’s a solid step toward running untrusted code securely on edge hardware, but it’s clearly not plug-and-play yet.
Elias: So that’s the reality of this work on "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges."
Priya: It shows that even in complex security domains, there are always performance and feature gaps we need to fill.
Nadia: Yeah, so as you listen, keep an eye on those updates regarding the ONNX runtime because that seems to be where this project is headed next.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel