OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design
cs.AR, cs.CR, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)
Code: https://github.com/intel/hexl-fpga
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Privacy in Deep Learning: A Survey
- Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network Inference
- Gazelle: A Low Latency Framework for Secure Neural Network Inference
- PrivCirNet: Efficient Private Inference via Block Circulant Transformation
- NeuJeans: Private Neural Network Inference with Joint Optimization of Convolution and FHE Bootstrapping
- AESPA: Accuracy Preserving Low-degree Polynomial Activation for Fast Private Inference
- Orion: A Fully Homomorphic Encryption Framework for Deep Learning
- CryptoGCN: Fast and Scalable Homomorphically Encrypted Graph Convolutional Network Inference
- Trinity: A General Purpose FHE Accelerator
- Osiris: A Systolic Approach to Accelerating Fully Homomorphic Encryption
- FAB: An FPGA-based Accelerator for Bootstrappable Fully Homomorphic Encryption
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU
- CryptGPU: Fast Privacy-Preserving Machine Learning on the GPU
- Very Deep Convolutional Networks for Large-Scale Image Recognition
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4