Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges

arXiv:2610.12050 · cs.CR, cs.SE · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Protecting CPU AI On Edge TEEs".

Elias: The gist: This work presents a solution that allows for the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we're looking at this paper, "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges." It’s talking about how to actually run existing AI models on edge hardware without giving the intellectual property away when someone gets root access.

Elias: Yeah, it tackles the problem that TEEs like Arm TrustZone are great for security but they make running unaltered applications inside them pretty hard without a ton of rewriting stuff. This paper proposes using WebAssembly to make that possible in OP-TEE.

Priya: What I wonder is how much of this protection actually translates into real confidentiality when you're dealing with these kinds of edge devices where things are always moving around.

Nadia: Exactly, Priya, and the paper suggests a specific way to do it: they compile those AI models into WebAssembly binaries first. Then, they encrypt those binaries and lock the key away in a physical fuse on the device itself so only the Trusted Application can unlock and run them.

Elias: And that whole setup relies on using OP-TEE with its WebAssembly Micro Runtime, or WAMR, to handle translating those WebAssembly instructions into what the TEE can actually execute natively.

Priya: So, if we’re talking about the real world, this means they're taking models that are already built and just wrapping them in this secure format instead of having to rebuild the entire application for the TEE.

Nadia: That's right. The paper is focused on answering a few specific research questions, like how WebAssembly can recompile existing AI models for Arm TrustZone without needing any changes to the original model itself.

Elias: And they also look at performance overhead, which is something everyone cares about when you're dealing with edge devices where resources are tight. They specifically ask what the performance overhead is when running these WebAssembly applications inside Arm TrustZone compared to native code.

Priya: I mean, for someone just driving or cooking, how does that translate? Is it a noticeable slowdown in startup time or during actual use?

Nadia: Well, the evaluation shows they found an additional overhead of twenty-two percent when comparing their solution to an application that was manually ported directly into OP-TEE.

Elias: That’s a significant number, so you can see it's not free. But they also noted that for inference latency, specifically for a MobileNetV2 model, they saw a mean slowdown of twenty-one point seven microseconds, which is about five point six one percent slower than running it in the WAMR environment inside OP-TEE.

Priya: So the numbers show it adds some friction to the operation compared to a perfectly ported application. What about the security side? How robust is this encryption mechanism they’re suggesting?

Nadia: The security model is pretty strict; model confidentiality is maintained because you can’t get the decryption key, which they state only exists within the Secure World and can be read by the WAMR Trusted Application.

Title and authors: Elias: They also talk about platform confidentiality, ensuring that the OP-TEE OS verifies any Trusted Application being loaded against a public key bundled with it to stop anyone from loading a custom application in there.

Priya: That addresses the risk of someone loading their own malicious code into the TEE, which is a big win for protecting that local AI model IP.

Nadia: But then they also flagged some clear limitations. They mentioned there are restrictions because of the limited support for current workloads and specifically a limited set of supported ONNX operators, which means you can't just run any random model format.

Elias: And another practical challenge they point out is that the solution currently lacks GPU acceleration, which definitely limits how fast these AI models can actually perform their tasks on the hardware.

Priya: So what does this mean for someone listening? It means you get a secure way to keep your AI model safe from theft when it’s running on your device, but you have to be very careful about which specific models you use and how much speed you expect.

Nadia: That's the summary of the trade-off: security and portability in exchange for some overhead and feature limitations right now. We need to see more work on expanding that AI model support, maybe by switching to an ONNX runtime, because that’s where they feel the biggest opportunity is left.

Elias: If we look ahead, they suggest future work will involve advancing the associated WASI neural network proposal and switching to a proper ONNX runtime. That would hopefully open up those performance gaps they are currently seeing in inference speed.

Priya: I’m just hoping that as the ecosystem matures, these kinds of solutions become more practical for everyday edge computing needs rather than just high-end research prototypes.

Nadia: That’s the direction we need to watch. So, to wrap up on this paper, "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges," it shows a way to execute unaltered AI models in OP-TEE for Arm TrustZone using encrypted WebAssembly binaries.

Elias: It’s a feasible approach that addresses the difficulty of running unaltered applications inside TEEs by leveraging WAMR and a custom distribution method where the encryption key is stored in device fuses.

Priya: And while it adds a twenty-two percent overhead compared to native ports, it still offers model confidentiality against root access attacks when dealing with ONNX formatted models.

Nadia: The real constraint is that the current implementation is restricted by the limited set of supported ONNX operators and the lack of GPU acceleration, so future work needs to focus on broadening those capabilities.

Elias: That’s where they are heading, aiming for an ONNX runtime and advancing that WASI neural network proposal to fix those performance issues.

Priya: So, for anyone listening who cares about the practical application, this paper gives us a concrete roadmap of what needs to be done next to make this technology truly ready for widespread edge deployment.

The paper's summary: Nadia: So, we’re looking at this paper that shows how you can run existing AI models on edge hardware without giving away the intellectual property when someone gets root access.

Elias: It tackles the problem that TEEs like Arm TrustZone are great for security but they make running unaltered applications inside them pretty hard without a ton of rewriting stuff.

Priya: What I wonder is how much of this protection actually translates into real confidentiality when you're dealing with these kinds of edge devices where things are always moving around.

Nadia: Exactly, Priya, and the paper suggests a specific way to do it: they compile those AI models into WebAssembly binaries first. Then, they encrypt those binaries and lock the key away in a physical fuse on the device itself so only the Trusted Application can unlock and run them.

Elias: And that whole setup relies on using OP-TEE with its WebAssembly Micro Runtime, or WAMR, to handle translating those WebAssembly instructions into what the TEE can actually execute natively.

Priya: So, if we’re talking about the real world, this means they're taking models that are already built and just wrapping them in this secure format instead of having to rebuild the entire application for the TEE.

Nadia: That's right. The paper is focused on answering a few specific research questions, like how WebAssembly can recompile existing AI models for Arm TrustZone without needing any changes to the original model itself.

Elias: And they also look at performance overhead, which is something everyone cares about when you're dealing with edge devices where resources are tight. They specifically ask what the performance overhead is when running these WebAssembly applications inside Arm TrustZone compared to native code.

Priya: I mean, for someone just driving or cooking, how does that translate? Is it a noticeable slowdown in startup time or during actual use?

Nadia: Well, the evaluation shows they found an additional overhead of twenty-two percent when comparing their solution to an application that was manually ported directly into OP-TEE.

Elias: That’s a significant number, so you can see it's not free. But they also noted that for inference latency, specifically for a MobileNetV2 model, they saw a mean slowdown of twenty-one point seven microseconds, which is about five point six one percent slower than running it in the WAMR environment inside OP-TEE.

Priya: So the numbers show it adds some friction to the operation compared to a perfectly ported application. What about the security side? How robust is this encryption mechanism they’re suggesting?

Nadia: The security model is pretty strict; model confidentiality is maintained because you can’t get the decryption key, which they state only exists within the Secure World and can be read by the WAMR Trusted Application.

Elias: They also talk about platform confidentiality, ensuring that the OP-TEE OS verifies any Trusted Application being loaded against a public key bundled with it to stop anyone from loading a custom application in there.

Priya: That addresses the risk of someone loading their own malicious code into the TEE, which is a big win for protecting that local AI model IP.

Nadia: But then they also flagged some clear limitations. They mentioned there are restrictions because of the limited support for current workloads and specifically a limited set of supported ONNX operators, which means you can't just run any random model format.

Elias: And another practical challenge they point out is that the solution currently lacks GPU acceleration, which definitely limits how fast these AI models can actually perform their tasks on the hardware.

Priya: So what does this mean for someone listening? It means you get a secure way to keep your AI model safe from theft when it’s running on your device, but you have to be very careful about which specific models you use and how much speed you expect.

Nadia: That's the summary of the trade-off: security and portability in exchange for some overhead and feature limitations right now. We need to see more work on expanding that AI model support, maybe by switching to an ONNX runtime, because that’s where they feel the biggest opportunity is left.

The paper's improvements: Nadia: So, we’re looking at what the authors suggest as improvements for this AI model execution setup on edge TEEs using WebAssembly.

Elias: They are basically saying they need to move past the current limitations because right now it’s too restrictive for real-world AI.

Priya: What exactly are they suggesting changes that would make this more useful for people actually building these systems?

Nadia: They're focusing on broadening the support beyond just a limited set of ONNX operators, which means they need to get better at handling different types of math operations in the AI models.

Elias: Yeah, and they point out that since it’s currently missing GPU acceleration, that’s a big thing because it limits how fast these models can actually perform their tasks on the hardware.

Priya: So if you want to run a faster model, this setup isn't going to give you that speed right now because of the hardware side?

Nadia: Exactly. The paper suggests future work needs to involve switching to a proper ONNX runtime, which should help with that performance issue.

Elias: And they also mention advancing the associated WASI neural network proposal, which is supposed to open up more ways for AI workloads to run securely inside the TEE environment.

Priya: So what does this mean for someone who just listens to the show? It means that for now, you can only use a narrow range of models, and you’re looking at some speed bumps.

Nadia: That’s the reality. They are calling out that the solution is currently restricted by memory management functions not part of the GlobalPlatform API, which they say OP-TEE Core provides as extensions instead.

Elias: So even if you get better operator support and a runtime, you still have to deal with those extra layers of complexity from the TEE core itself.

Priya: That sounds like it puts a lot of work on the developers who are trying to make these things usable for everyone, not just high-end research prototypes.

Nadia: Right. The authors are basically laying out a roadmap where they need to focus on that ONNX runtime and the WASI proposal if they want this thing to actually scale up beyond what it is today.

Conclusion: Tom: So we’re wrapping up on this paper, "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges," which shows how you can run existing AI models on edge hardware without giving away the intellectual property when someone gets root access.

Nadia: It basically confirms that WebAssembly combined with OP-TEE is a way to get unaltered AI models running inside secure environments like Arm TrustZone.

Elias: The main point here is using that custom distribution method where the encryption key stays in device fuses so only the trusted application can unlock and execute the model.

Priya: So what does this actually change for someone who just listens to the show? It means they can protect their local AI models from being stolen when they're running on their phone or device without needing to completely rewrite everything.

Nadia: Right. But we have to remember those performance trade-offs, like that twenty-two percent overhead compared to a manually ported application.

Elias: And the inference latency still shows a slowdown for models like MobileNetV2, which is about five point six one percent slower in the TEE WAMR environment.

Priya: I just want to focus on what the data actually shows: it’s secure against root access, but it’s not perfectly fast either right now.

Nadia: That's the constraint they are flagging—the limited support for specific ONNX operators and the lack of GPU acceleration are big roadblocks for what this can do today.

Elias: They are setting a clear path forward by saying future work has to involve switching to an ONNX runtime and advancing that WASI neural network proposal.

Priya: So, the implication is that this technology is feasible for protecting AI IP, but we need more development on the toolset to make it practical for real-world performance.

Nadia: Exactly. It’s a solid step toward running untrusted code securely on edge hardware, but it’s clearly not plug-and-play yet.

Elias: So that’s the reality of this work on "Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges."

Priya: It shows that even in complex security domains, there are always performance and feature gaps we need to fill.

Nadia: Yeah, so as you listen, keep an eye on those updates regarding the ONNX runtime because that seems to be where this project is headed next.

Friedrich Vandenberghe, Lachlan Gunn, Bruno Volckaert, Merlijn Sebrechts

IDLab, Department of Information Technology Ghent University - imec, Belgium · Aalto University, Finland

cs.CR, cs.SE

Submitted: 2026-10-08

Updated: 2026-10-08

Comments: 8 pages, 5 figures

Code: https://github.com/TephrocactusMYC/TREO

License: http://creativecommons.org/licenses/by-sa/4.0/

The gist: The gist: This work presents a solution that allows for the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone.

Key concepts

Arm TrustZone
This is a hardware security extension that divides a device's processor into two worlds: the Normal World and the Secure World. The Secure World has special access to sensitive resources, ensuring that critical operations and data remain protected from potentially compromised applications running in the Normal World.
WebAssembly (Wasm)
Wasm is a portable compilation target for existing applications. It allows developers to compile AI models into this format, which can then be executed securely inside the TEE. This portability solves the problem of having to heavily modify AI software just to run it in a constrained TEE environment.
WAMR (WebAssembly Micro Runtime)
WAMR is the specific runtime used within OP-TEE to execute WebAssembly code. It acts as a translator, taking Wasm instructions and converting them into native operations that the OP-TEE system can understand and run securely on the hardware.

Terminology

Summary

The gist: This work presents a solution that allows for the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone.

How it works

The proposed solution enables the execution of unaltered AI models inside an edge Trusted Execution Environment (TEE) such as OP-TEE for Arm TrustZone by leveraging WebAssembly and a custom distribution methodology. This approach addresses the challenge that TEEs like Arm TrustZone provide a constrained environment, making it difficult to run unaltered applications inside them without significant porting effort. The system utilizes the WebAssembly Micro Runtime (WAMR) to execute these workloads within OP-TEE.

The distribution methodology involves encrypting the WebAssembly binary and placing the encryption key in one of the device’s fuses. This ensures that only the WAMR Trusted Application (TA) in OP-TEE can decrypt and execute the AI model. The architecture involves a trusted AI model distributor which trains, signs, and encrypts an AI model with a model key.

Key Components and Technologies

The solution integrates several technologies to achieve secure execution of encrypted WebAssembly code using Arm TrustZone.

  1. Arm TrustZone provides the fundamental division of the system into ‘Normal World’ and ‘Secure World’ modes, with Trusted Applications (TA) having access to a secure physical address space.

  2. OP-TEE is an open-source TEE OS that implements an API for communication between TEE and TA, facilitating the transition from the Normal World to the TA.

  3. WebAssembly (Wasm) serves as a portable compilation target, allowing existing applications to be compiled into a format executable inside the TEE.

  4. The WebAssembly Micro Runtime (WAMR) is used as the runtime to translate Wasm instructions into native ones within OP-TEE.

  5. The AI models are distributed in the Open Neural Network Exchange (ONNX) format, which allows for easy porting of existing applications to run inside the TEE.

Evaluation and Performance

A thorough evaluation of the solution shows that it incurs an additional overhead of 22% in comparison to an application manually ported to OP-TEE. Furthermore, this method facilitates the execution of unaltered AI models with an additional inference latency of 6% without significant porting effort.

The performance evaluation measures both startup and runtime latency. Startup latency analysis shows that the decryption time scales linearly with the binary size, as demonstrated by execution times ranging from 1 to 10 MiB.

In terms of inference latency for a MobileNetV2 model, switching from the REE WAMR environment to TEE WAMR results in a mean slowdown of 21.7µs (+5.61%). The greatest overhead is attributed to making use of WAMR (+23.6%), not the switch from REE to TEE (+2.62%).

Security Guarantees and Limitations

The security model focuses on achieving two main objectives: Model confidentiality and Platform confidentiality.

  1. Model confidentiality is maintained because the adversary cannot access the decryption key, which is derived with the NIST-SP 800-108 KDF from an AES key only readable from within the Secure World. The use of authenticated encryption prevents chosen-ciphertext attacks against model weights.

  2. Platform confidentiality is preserved because the OP-TEE OS verifies the signature of any loaded TA against an RSA public key bundled with the OP-TEE OS, preventing adversaries from loading a custom-designed TA.

However, limitations exist in the current implementation regarding AI workloads. Specifically, there is a restriction due to the limited support for current workloads and the limited set of supported ONNX operators. Additionally, the solution lacks GPU acceleration, which limits the performance of the AI model. Future work will be necessary to broaden this support by switching to an ONNX runtime and advancing the associated WASI neural network proposal. The solution is also restricted by memory management functions not part of the GlobalPlatform API, which are provided as extensions by OP-TEE Core.

Conclusion

This work enables executing unmodified AI models in a manner such that an attacker with root access cannot exfiltrate them, and thus the intellectual property of the AI models remains safe. While this is a feasible approach, it is heavily restricted by the limited set of supported ONNX operators and the limited support for machine learning workloads inside WASI. The solution incurs a 22% slowdown when compared to a native execution inside OP-TEE, which is 0.7% faster than the 2021 WAMR version used in WaTZ. Future work will be necessary to broaden the supported set of AI models, which will require switching to an ONNX runtime and advancing the associated WASI neural network proposal. The performance-wise, our solution with a recent version of WAMR incurs a 22% slowdown when compared to a native execution inside OP-TEE, which is 0.7% faster than the 2021 WAMR version used in WaTZ. The inference latency of MobileNetV2 shows a 5.61% slowdown when compared to an execution of WAMR inside OP-TEE. The solution is heavily restricted by the limited set of supported ONNX operators and the limited support for machine learning workloads inside WASI. The solution lacks GPU acceleration, which limits the performance of the AI model. Future work will be necessary to broaden this support by switching to an ONNX runtime and advancing the associated WASI neural network proposal. The solution is also restricted by memory management functions not part of the GlobalPlatform API, which are provided as extensions by OP-TEE Core.

ACKNOWLEDGEMENT

This work received funding from IMEC AAA PSX-PSCEC and the European Union under Grant Agreement number 101225859. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. The European Union cannot be held responsible for them.

REFERENCES

[1] Arm, “Learn the architecture - TrustZone for AArch64,” Oct. 2024. [Online]. Available: https://developer.arm.com/documentation/102418/latest/

[2] ——, “Arm security technology building a secure system using TrustZone technology,” Apr. 2009. [Online]. Available: https://developer.arm.com/documentation/PRD29-GENC-009492/c

[3] TrustedFirmware.org. [Online]. Available: https://optee.readthedocs.io/en/latest/architecture/libraries.html

[4] W3C, “WebAssembly.” [Online]. Available: https://webassembly.org/

[5] Bytecode Alliance, “WebAssembly Micro Runtime.” [Online]. Available: https://github.com/bytecodealliance/wasm-micro-runtime

[6] onnx, “ONNX.” [Online]. Available: https://onnx.ai/

[7] F. Vandenberghe, “idlab-discover/watz.” [Online]. Available: https://github.com/idlab-discover/watz

[8] NVIDIA, “NVIDIA Jetson Linux developer guide - secure boot.” [Online]. Available: https://docs.nvidia.com/jetson/archives/r36.4.3/DeveloperGuide/SD/Security/SecureBoot.html

[9] NVIDIA, “OP-TEE in Jetson Linux.” [Online]. Available: https://docs.nvidia.com/jetson/archives/r38.4-devguide-sd-security-optee.html

[10] Linaro, “OP-TEE.” [Online]. Available: https://www.trustedfirmware.org/projects/op-tee/

[11] GlobalPlatform, “GlobalPlatform.” [Online]. Available: https://globalplatform.org/

[12] TrustedFirmware.org, “OP-TEE pseudo trusted applications.” [Online]. Available: https://optee.readthedocs.io/en/latest/architecture/trusted-applications

[13] ——, “OP-TEE frequently asked questions.” [Online]. Available: https://optee.readthedocs.io/en/latest/faq/faq.html

[14] W3C, “WebAssembly.” [Online]. Available: https://webassembly.org/

[15] appcypher, “Awesome WebAssembly languages.” [Online]. Available: https://github.com/appcypher/awesome-wasm-langs

[16] D. Gohman, Yosh, D. Bakker, L. Clark, P. Hickey, A. Crichton, B. Hayes, L.

Improvements for AI systems

  1. The improved system can execute unaltered AI models by leveraging WebAssembly (Wasm) compiled for execution inside OP-TEE for Arm TrustZone, addressing the challenge of running existing AI software inside a TEE without significant porting effort.

  2. The model distributor can encrypt the WebAssembly binary and place the encryption key in device fuses, ensuring only the WAMR Trusted Application (TA) in OP-TEE can decrypt and execute the AI model.

  3. The system will maintain Model confidentiality by ensuring that even an adversary with root access cannot exfiltrate weights because they cannot access the decryption key, which is derived within the Secure World.

  4. The platform supports execution of models compiled to the Open Neural Network Exchange (ONNX) format, allowing for easy porting of existing applications, thus fulfilling the promise of protecting important AI model IP against adversaries with root access.

  5. The system demonstrates a performance profile where it incurs an additional overhead of 22% in comparison to an application manually ported to OP-TEE, while also facilitating the execution of unaltered AI models with an additional inference latency of 6%.

Abstract

AI models on edge hardware contain important intellectual property (IP), but an adversary can steal it when they achieve root access. Trusted Execution Environments (TEE) like Arm TrustZone protect against these Operating System (OS) level attacks. However, they are challenging to use. More precisely, it is difficult to run unaltered applications inside a TEE. This work presents a solution that allows the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone. Additionally, this work provides an AI model distributor that encrypts the WebAssembly binary and places the encryption key in one of the device's fuses. This way, only the WAMR Trusted Application (TA) in OP-TEE can decrypt and execute the AI model. A thorough evaluation of our solution shows that it incurs an additional overhead of 22% in comparison to an application manually ported to OP-TEE, while also facilitating the execution of unaltered AI models with an additional inference latency of 6% without significant porting effort. Overall, there is a pressing need to safeguard the IP of AI models and this work shows that there is a real promise in WebAssembly, but there remain some practical challenges.

Sources

Related papers