Paper 2026/1862
HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation
Abstract
Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrapping operation is required to continue. Every nonlinearity must therefore be approximated by an iterative method; each iteration increasing the number of multiplications. A higher iteration count buys precision but exhausts the available depth more frequently and thus triggers more bootstraps, which dominate latency. We introduce Homomorphic Encryption-Aware Training (HEAT), a fine-tuning method that makes the per-nonlinearity iteration counts learnable, enabling them and the model weights to co-adapt during training. HEAT optimizes iterations with respect to the task objective, allowing the model to adapt to approximation errors encountered during inference without architectural changes or retraining from scratch. We further relate iteration count to quantization bit width and bound, at fixed weights, the gap between our objective and quantization-aware training. On encrypted GPT-2 decoding, HEAT reduces iterations by $3.1\times$, bootstraps by $1.6\times$, and end-to-end latency by $1.4\times$, while improving decode agreement over the calibrated encrypted baseline.
Note: 9 pages, original extended abstract (v1, 4 pages) accepted at "Privacy in the Era of Large Opaque Models: Theoretical, Legal, and Practical Perspectives" and "Beyond Private Training: The New Landscape of AI Privacy"
Metadata
- Available format(s)
-
PDF
- Category
- Applications
- Publication info
- Preprint.
- Keywords
- Fully Homomorphic EncryptionMachine LearningFine-tuningCKKSLanguage ModelsLLMsDiscrete OptimizationGPU
- Contact author(s)
-
zirilli @ di uniroma1 it
marincione @ di uniroma1 it
evgenions @ gmu edu
ateniese @ gmu edu
rodola @ di uniroma1 it - History
- 2026-09-30: revised
- 2026-09-02: received
- See all versions
- Short URL
- https://ia.cr/2026/1862
- License
-
CC BY-NC-SA
BibTeX
@misc{cryptoeprint:2026/1862,
author = {Alessandro Zirilli and Davide Marincione and Evgenios M. Kornaropoulos and Giuseppe Ateniese and Emanuele Rodolà},
title = {{HEAT}: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation},
howpublished = {Cryptology {ePrint} Archive, Paper 2026/1862},
year = {2026},
url = {https://eprint.iacr.org/2026/1862}
}