On the Sensitivity to Errors in Homomorphic Computing: Single Transient Bit-flip Client-side Error Characterization

arXiv:2608.11155 · cs.AR, cs.CR · Submitted 2026-08-11 · Read on arXiv

Matías Mazzanti, Vattana Chan, Karthik Swaminathan, Augusto Vega, Esteban Mocskos, Radha Venkatagiri

University of Buenos Aires · IBM T. J. Watson Research Center · Georgetown University

cs.AR, cs.CR

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 3 pages, 3 figures

Code: https://github.com/Microsoft/SEAL

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper analyzes the sensitivity of Homomorphic Encryption (HE) to bit-level faults, focusing on the CKKS (Cheon–Kim–Kim–Song) scheme, which is widely used for approximate arithmetic in AI

Terminology

Summary

This paper analyzes the sensitivity of Homomorphic Encryption (HE) to bit-level faults, focusing on the CKKS (Cheon–Kim–Kim–Song) scheme, which is widely used for approximate arithmetic in AI and machine learning workloads. The authors identify homomorphic multiplication as the most error-sensitive operation in practical HE pipelines and characterize how faults propagate and amplify through it, exposing a critical robustness vulnerability and motivating the need for more resilient HE deployments.

The paper presents three main findings: 1) error resilience of FHE on client-side can be categorized into two patterns: Addition pattern and Multiplication pattern across Vanilla and RNS configurations; 2) Multiplication pattern dominates addition pattern when a multiplication is performed on the server; and 3) the two resilience patterns observed are parametric and data independent.

The authors use a custom C++ implementation of CKKS (C-CKKS) that supports native RNS, NTT, combined RNS+NTT arithmetic, and 64-bit coefficient representations. They refer to the configuration where both RNS and NTT optimizations are disabled as Vanilla CKKS. Fault injection experiments are conducted using the open-source LLTFI framework. They adopt a single-bit error model, where each bit of every polynomial coefficient (in both the plaintext and the ciphertext) is flipped in sequence, and after each bit error injection, they execute the entire HE pipeline and compare the recovered data after decoding against the original data using the Maximum Relative Error Percentage (MREP).

For the Addition pattern, the authors observe that in Vanilla CKKS, bit-flips affecting values below the scaling factor (∆) or exceeding the modulus (Q) resilience lead to masked plaintext and both ciphertext components, c0 and c1. Bit-flips occurring on gap coefficients or on the N/2-th coefficient consistently produce fully masked effects in the plaintext and in c0, but not in c1. This behavior arises because gaps are present in the plaintext and are therefore encoded into c0, whereas c1 corresponds to a sampled secret-key-dependent component. A closely related pattern is observed in RNS CKKS; however, individual coefficients do not exhibit scaling-factor (∆) or modulus (Q) resilience, as each coefficient is represented using multiple limbs.

For the Multiplication pattern, an error-resilience profile closely resembling that of addition is observed in both Vanilla CKKS and RNS CKKS. However, in Vanilla CKKS, gap and N/2 coefficients remain resilient only up to a certain error threshold, whereas in RNS CKKS these coefficients exhibit no resilience and are completely corrupted. In contrast, CKKS configurations employing NTT and RNS+NTT arithmetic are fully susceptible to bit-flips, as any single bit-flip ultimately results in catastrophic decryption and decoding failures. Consequently, no distinct addition or multiplication error patterns can be identified under these configurations.

The authors also observe that whenever a multiplication is performed on the server, the overall resilience profile consistently follows the multiplication pattern, indicating that multiplication pattern dominates the addition pattern. Furthermore, the observed resilience heuristics are intrinsic to the scheme itself, rendering them independent of both parameters and data.

The paper concludes that errors occurring at different stages of CKKS execution—encoding, encryption, decryption, and decoding—can lead to three outcomes: either the HE execution fails and the error is explicitly detected at decoding, the execution completes successfully while the error propagates across stages and manifests as silent data corruption (SDC), or the error may be masked and have no observable effect on the output. The severity of SDC in HE applications stems from the error-centric nature of HE schemes themselves, as in the encrypted domain, data is inherently noisy by construction, and additional errors induced by faulty hardware or software can become indistinguishable from the intrinsic HE noise, making detection extremely challenging.

Improvements for AI systems

Improvements to AI Systems:

  1. Fault-Aware HE Compiler/Runtime Scheduler: Integrate the identified Addition and Multiplication resilience patterns into an AI inference compiler. The system can automatically reorder or fuse operations (e.g., batching additions before multiplications) to minimize exposure to bit-flips, or insert lightweight checksums only on multiplication-heavy layers where corruption is most likely.

  2. Adaptive Error Detection for HE-Based ML Inference: Build an AI model that learns to predict which ciphertext coefficients are at risk (e.g., gap or N/2 coefficients in Vanilla/RNS CKKS) based on the operation history. This predictor can trigger selective re-encryption or redundant computation only for high-risk operations, reducing overhead while catching silent data corruption (SDC) before it propagates to final predictions.

  3. Resilience-Aware Neural Architecture Search (NAS): Use the paper’s finding that multiplication dominates error patterns to guide NAS for HE-compatible models. The improved system can penalize architectures with deep multiplication chains (e.g., many consecutive multiplications) and favor those with more additions or parameter-efficient multiplications, yielding models that are inherently more robust to hardware faults without sacrificing accuracy.

  4. Fault-Injection-Driven Training Regularization: Simulate the observed bit-flip patterns (e.g., flipping gap coefficients or limbs in RNS) during training of a small meta-model that predicts post-decryption error. This meta-model can then be used to fine-tune the main AI model to be more tolerant to such faults, effectively making the model’s weights and activations less sensitive to specific coefficient corruptions.

  5. Runtime Health Monitoring for HE Accelerators: Deploy a lightweight AI anomaly detector that monitors MREP trends across consecutive HE operations. Since the paper shows resilience is parametric and data-independent, the detector can learn the expected error profile for a given configuration and flag deviations that indicate active bit-flips, enabling early system-level intervention (e.g., task migration or key refresh).

  6. Configuration Selection Advisor: An AI-based advisor that, given a target AI workload (e.g., transformer vs. CNN) and hardware fault rate, recommends the optimal CKKS configuration (Vanilla, RNS, NTT, or RNS+NTT) and parameter choices (scaling factor, modulus size) to balance performance and robustness, using the paper’s finding that NTT-based configs are fully susceptible and should be avoided in high-fault environments.

What the improved AI system can do:

  • Run HE-based inference with significantly lower risk of silent data corruption, especially in multiplication-heavy models like deep neural networks.

  • Automatically adapt its execution strategy (operation ordering, redundancy, and configuration) in real time based on observed fault patterns.

  • Provide predictable and verifiable robustness guarantees for AI workloads in untrusted or faulty hardware environments (e.g., edge devices, cloud servers with aging silicon).

  • Reduce the overhead of fault mitigation by focusing resources only on the most vulnerable operations (multiplications) and coefficients (gap/N/2), rather than blanket redundancy.

Sources

Related papers