ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation
cs.CR
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 15 pages. Accepted at ACM CCS 2026
Code: https://github.com/tuneinsight/lattigo
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- Trinity: A General Purpose FHE Accelerator
- Large Language Model Agent in Financial Trading: A Survey
- Orion: A Fully Homomorphic Encryption Framework for Deep Learning
- The Llama 3 Herd of Models
- NeuJeans: Private Neural Network Inference with Joint Optimization of Convolution and FHE Bootstrapping
- Gazelle: A Low Latency Framework for Secure Neural Network Inference
- MPCFormer: fast, performant and private Transformer inference with MPC
- Pointer Sentinel Mixture Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Efficient Streaming Language Models with Attention Sinks
- HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
- Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network Inference
- PrivCirNet: Efficient Private Inference via Block Circulant Transformation
- PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
- Cachemir: Fully Homomorphic Encrypted Inference of Generative Large Language Model with KV Cache
- EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
- TinyLlama: An Open-Source Small Language Model
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs