MOSAIC: Masked Outsourcing of Secure AI Computations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MOSAIC: Masked Outsourcing of Secure AI Computations".
Jane: The paper was written by James Hsin-yu Chiang, Sheila Zingg, Kari Kostiainen and Srdjan Capkun from ETH Zurich.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Tom: So, we’ve established that "MOSAIC: Masked Outsourcing of Secure AI Computations" is a big deal, but what does the summary actually tell us? It explains this whole concept of securely outsourcing AI computations from a trusted but weak client to an untrusted but powerful server.
Jane: Essentially, the paper is addressing that huge challenge where the client holds both the data and the model—the intellectual property—but can’t handle it all computationally, so they need to offload it. But "MOSAIC" is designed so that neither, nor the untrusted accelerator, learns anything about what they are computing.
Lu: It's a clever workaround for trust; by keeping the trusted computing base or TCB small and outsourcing the bulk of the linear math, we bypass some of those traditional security walls.
Meng: I liked hearing that this scales up to modern workloads like 70B transformer inference, which is a huge leap from anything before.
Lalam: The summary makes it sound like a practical solution for enabling large-scale confidential AI in our data centers, which is incredibly important for the future of privacy.
Tom: It sounds like we are moving from theoretical concepts to seeing how this could actually work in a real environment, and that’s exactly what the next segment will cover.
Improvements: Tom: Moving beyond the general idea, what specific improvements does "MOSAIC: Masked Outsourcing of Secure AI Computations" offer over previous attempts at secure outsourcing? The paper points to some critical limitations in the state-of-the-art that "MOSAIC" seems to solve.
Jane: One major issue with older methods was that they just didn't scale up well enough, especially for large modern models, and "MOSAIC" achieves an optimal asymptotic client overhead of O((m+n)l). That's a massive improvement in efficiency.
Lu: The authors also tackled the problem of cryptographic proposals usually operating only in fixed-point domains, which is terrible for large floating-point LLM architectures. They’ve solved this with some sophisticated masking techniques.
Meng: I'm interested in how they manage the complexity; it sounds like the old solutions were O(mn epsilon l) client overhead, which is impractical for 72B models, but "MOSAIC" is far better.
Lalam: The introduction of noise and then carefully managing that error makes the computation robust and reliable, ensuring that accuracy stays very close to full precision in a way previous systems couldn't guarantee it.
Tom: It sounds like we’ are not just faster, but more reliable too, which is essential when balancing security and performance.
Conclusions: Tom: We've seen the technical improvements, but now let's wrap our discussion up by looking at the big picture—what does "MOSAIC: Masked Outsourcing of Secure AI Computations" really mean for the future?
Jane: Overall, it’s a practical path toward confidential AI where we can use untrusted hardware without having to expand our entire trusted computing base.
Lu: The paper proves that this approach works even on large 70B models, showing that the error accumulation is non-destructive and stays within practical limits.
Meng: I'm excited to see the implementation results, especially how it handles things like prefill and decoding in a real data center setting without bottlenecking.
Lalam: We can have AI that is both powerful for complex tasks and inherently private, which is a huge win for society as we move towards more sophisticated applications.
Tom: And finally, by recognizing the limitations of traditional TCBs, "MOSAIC: Masked Outsourcing of Secure AI Computations" provides a scalable alternative.
Lu: I think we have covered all the major points—the technical breakthroughs, the practical efficiency gains, and what this could mean for a huge range of future applications.
Meng: It really seems like that while my job is to make these systems run, "MOSAIC" gives us a roadmap that allows me to build at scale.
Lalam: My final thought is that this technology allows AI to grow in complexity without sacrificing the privacy of the people who are using it.
Tom: Thank you all for sharing your insights on this remarkable work by Chiang and Zingg et al., and we'll be back with another fascinating paper next time!
Conclusion: Tom: So, wrapping up our deep dive into "MOSAIC: Masked Outsourcing of Secure AI Computations," it really feels like we've covered ground that shifts how we think about AI infrastructure.
Jane: It’s wild to think that this paper tackles the fundamental tension between needing massive compute power and keeping proprietary data safe from the cloud provider itself.
Meng: Exactly. The practical hurdle here isn't just *if* we can outsource, but *how* we prove that outsourcing is secure enough for sensitive enterprise applications.
Lu: And what MOSAIC proposes—this masked outsourcing framework—is a serious leap because it moves beyond just encryption; it tackles the computation side itself.
Tom: It’s revolutionary in how it formalizes trust boundaries, essentially letting you use massive external resources without giving up control over the underlying data integrity.
Jane: I mean, for anyone who's ever worried about sending sensitive company algorithms to a third-party cloud service, this gives them a whole new level of confidence.
Meng: From an engineering standpoint, the ability to quantify and manage that trust boundary is what makes this commercially viable; it’s not just theory anymore.
Lu: I think the biggest implication is that it might accelerate the adoption of AI in highly regulated industries like finance or healthcare, where data sovereignty is paramount.
Tom: It's a massive enabler, Lu. It’s like handing a giant key to industries that have been hesitant to even start using advanced AI tools because of compliance fears.
Jane: And it makes the whole concept of "secure compute" much more accessible, which really democratizes access to powerful AI tools for smaller organizations too.
Meng: Absolutely, the cost reduction and risk mitigation combined means that companies that previously thought they couldn't afford secure AI are suddenly in the running.
Lalam: Considering its impact on culture, I believe "MOSAIC: Masked Outsourcing of Secure AI Computations" will fundamentally change how we perceive collaboration; it allows us to pool intellectual resources globally without sacrificing local control or privacy.
Tom: That's a powerful way to put it, Lalam. It’s building a global infrastructure for secure innovation!
Jane: Well, our time is flying, but what an amazing discussion this has been about truly secure and outsourced AI compute.
Meng: We've got a lot to think about regarding implementing this in real-world pipelines.
Lu: Keep reading up on the architectural implications of masked outsourcing; it’s genuinely groundbreaking work.
Lalam: Thanks for joining us today; we hope you all find these insights valuable as we move forward into the next topic.
ETH Zurich
cs.CR, cs.AI
Submitted: 2026-07-31
Updated: 2026-08-25
Code: https://github.com/jachiang/mosaic
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 4/100
The gist: The paper, "MOSAIC: Masked Outsourcing of Secure AI Computations," addresses the challenges and security implications associated with outsourcing Artificial Intelligence computations, particularly
Key concepts
- Masked Outsourcing
- This is the core concept of the paper, allowing secure outsourcing of AI computations. It ensures that neither the client nor the untrusted accelerator learns anything about what is being computed, providing a clever workaround for trust issues in large-scale AI systems.
- Trusted Computing Base (TCB)
- The TCB refers to the small set of components that must be trusted to ensure security. MOSAIC minimizes this TCB while allowing the bulk of the linear math computation to be outsourced, bypassing traditional security walls and improving scalability.
- Confidential AI
- This is a practical application goal. It allows for the use of untrusted hardware for complex AI tasks while maintaining data privacy and protecting intellectual property. This capability is seen as a major enabler for sensitive fields like finance.
Terminology
Summary
The paper, MOSAIC: Masked Outsourcing of Secure AI Computations,
addresses the challenges and security implications associated with outsourcing Artificial Intelligence computations, particularly for large language models (LLMs), to untrusted cloud environments over public networks.
Problem Context and Topology:
The core scenario investigated is the Remote outsourcing topology: a trusted client at the local site outsources a single forward pass per query to an untrusted cloud GPU over a public network (e.g. 10 ms RTT, 100 Gbps).
This setup simulates real-world industrial applications, such as Industrial diagnostics. Consider an industrial predictive maintenance scenario in which a company operates a fleet of equipment—pumps, compressors, motors, gearboxes—and wishes to outsource...
Methodology and Protocol:
The research utilizes a protocol involving Masked Outsourcing of Secure AI Computations.
The methodology involves simulating both network latency and computational runtimes. Specifically, the authors simulate network latency induced by our protocol for different network settings (Figures 15 and 16).
To evaluate the robustness of the outsourced computation, the study focuses on measuring error accumulation across model layers. This is quantified using metrics such as Relative 2 error
and Cosine similarity.
The analysis tests various quantization schemes, including 4-bit NF4,
INT8,
and INT16-rot,
under different levels of simulated protocol noise (sigma), ranging from sigma = 0 (noise-free) up to sigma = 2.0.
Experimental Scope and Models:
The evaluation is conducted across several state-of-the-art LLMs, including:
-
Qwen2.5-32B (64 model layers)
-
Qwen2.5-72B (80 model layers)
-
LLaMA-3-70B (80 layers)
-
DeepSeek-R1-Distill-LLaMA-70B (80 layers)
The analysis provides detailed metrics at the last two transformer model layers and after the final RMSNorm, presented in Table 7.
Key Findings and Results:
-
Error Accumulation Analysis: The paper presents per-layer error accumulation figures for multiple models (Figures 18 through 21). For instance, Figure 20 shows
Per-layer error accumulation: LLaMA-3-70B (80 layers).
The results demonstrate how the fidelity of the computation degrades as noise increases. For Qwen2.5-32B, at = 63, the relative 2 error with sigma = 1.0 is 0.0532. -
Quantization Performance: The study compares performance across quantization schemes, noting that for DeepSeek-R1-Distill-LLaMA-70B, the relative 2 error at = 78 with sigma = 1.0 is 0.5464.
-
Computational and Network Performance: The paper evaluates system performance under WAN conditions, detailing
System prefill and runtime latency under simulated WAN (10 ms RTT, 100 Gbps, G = 3 untrusted GPUs) for LLaMA-3-70B
(Figure 16). -
Model Sensitivity: The analysis highlights model sensitivity to noise and quantization. For example, Figure 20 notes that
INT8 baseline collapses at = 79 and goes off-scale,
suggesting specific vulnerabilities in certain models or schemes when subjected to protocol noise.
In summary, MOSAIC provides a comprehensive framework for analyzing the trade-offs between computational efficiency, network latency, and security when outsourcing AI inference to untrusted cloud resources by rigorously quantifying error accumulation across model layers using various quantization and noise profiles.
Improvements for AI systems
The research presented details a comprehensive framework (MOSAIC) for secure, reliable, and efficient outsourcing of Large Language Model (LLM) inference. The core improvements should focus on resilience engineering and resource-aware orchestration when deploying LLMs in hybrid cloud environments.
-
Improvement: Integrate a dynamic, adaptive error mitigation layer into the model's architecture, specifically targeting the residual stream output (r l) at each transformer layer. This module must dynamically select between different numerical representations (e.g., BF16, INT16-rot, or quantized formats like 4-bit NF4) based on the measured cumulative error rate and the expected noise profile (sigma).
-
Mechanism: The system should continuously monitor metrics such as the relative 2 error and cosine similarity across successive layers. If the error accumulation exceeds a predefined threshold (e.g., 10-3 for INT16-rot), the system must automatically trigger a fallback or weight adjustment mechanism, such as temporarily switching to a higher precision format or applying masked redundancy coding only to the most sensitive layers identified by prior analysis (e.g., the final RMSNorm).
-
What it enables: The improved system can maintain high inference accuracy (>99% similarity) even when operating under highly noisy, low-precision, or resource-constrained outsourced environments, preventing catastrophic model collapse (INT8 failure observed at =79).
-
Improvement: Develop a sophisticated scheduling and routing layer that treats network topology (Local Site Cloud GPU) not just as a latency variable, but as an integral part of the computational cost function. This layer must dynamically partition the LLM inference task (e.g., token generation) into segments optimized for the current WAN conditions (RTT, Bandwidth).
-
Mechanism: Instead of assuming a fixed single-pass architecture, the system should implement Adaptive Task Chunking. For high-latency links (>10 ms RTT), it prioritizes outsourcing only compute-intensive, non-sequential blocks (e.g., attention key/value projections) while keeping critical, low-latency components (like initial token embedding or final output decoding) local. It must quantify the trade-off between computational overhead on the client vs. network latency penalty (Time total = Time client compute + sum (Latency + DataSize over Bandwidth)).
-
What it enables: The system can execute complex, multi-step LLM workflows (e.g., RAG pipeline components) reliably over unreliable or high-latency public networks, maximizing throughput by minimizing the impact of network bottlenecks and intelligently masking data transfers.
-
Improvement: Formalize the secure outsourcing protocol by integrating a verifiable computational attestation mechanism alongside the data transfer. This moves beyond simple encryption to include proof that the outsourced operation was executed correctly, even if the hardware is untrusted.
-
Mechanism: The client must mandate periodic Intermediate State Verification Checks. After an outsourced forward pass, instead of only receiving the result logits, the cloud GPU must also return a cryptographic proof (e.g., a zero-knowledge proof or a simple checksum/hash) confirming that the arithmetic operations were performed according to the specified protocol and precision (INT16-rot). This creates a robust feedback loop that detects silent data corruption or malicious manipulation on the remote hardware.
-
What it enables: The system provides provable security assurance for outsourced AI computations, allowing deployment in highly regulated or mission-critical industries (like industrial diagnostics) where simple encryption is insufficient to guarantee computational integrity.
Abstract
We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server must learn neither. We present MOSAIC, whose core is a novel matrix-multiplication masking protocol that scales to far larger matrices than prior work, enabling the safe outsourcing of modern workloads such as large transformer inference. By introducing small amounts of noise to the multiplication result and thereby relaxing correctness, MOSAIC achieves optimal asymptotic client overhead and concrete runtimes orders of magnitude faster than prior work. Its security reduces to the decisional LWE and LPN assumptions. Because this noise accumulates across the many layers of a transformer, a key technical challenge is bounding error growth; MOSAIC addresses this with an error-scaling mechanism based on random Hadamard rotations. On large 70B transformer models, MOSAIC's perplexity is comparable to popular quantization approaches and even matches full-precision BF16 inference on HumanEval. Finally, we present an end-to-end implementation showing how ideas like MOSAIC can promise a path towards large-scale confidential AI in modern data centers. Non-confidential inference is already distributed across phase (prefill/decode), layer, and time to maximize utilization of heterogeneous hardware, using RDMA-like networking to move activations, cached KV values, and weights across nodes. MOSAIC enables scaling of confidential compute by keeping the trusted computing base (TCB) small and outsourcing the bulk of the AI computation to untrusted accelerators.
Sources
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- Practical Secure Delegated Linear Algebra with Trapdoored Matrices
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- QLoRA: Efficient Finetuning of Quantized LLMs
- Blueprint, Bootstrap, and Bridge: A Security Look at NVIDIA GPU Confidential Computing
- SpinQuant: LLM quantization with learned rotations
- The Uniqueness of LLaMA3-70B Series with Per-Channel Quantization
- Lipschitz regularity of deep neural networks: analysis and efficient estimation
- Multi-stage Flow Scheduling for LLM Serving
- Slalom: Fast, Verifiable and Private Execution of Neural Networks in Trusted Hardware
- PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- Power-Softmax: Towards Secure LLM Inference over Encrypted Data
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs