Paper 2026/2356
BOLT-FHE: An Efficient Unified Framework for GPU-based TFHE Bootstrapping via On-Chip Local Tiling Strategies
Abstract
Bootstrapping is the main performance bottleneck in bitwise Fully Homomorphic Encryption (FHE), and practical acceleration requires careful orchestration of the blind rotation and external product chain under GPU resource constraints. This paper presents BOLT-FHE, a GPU bootstrapping framework that emphasizes block-local execution, on-chip tiling, and a unified MegaKernel supporting both gadget decomposition and modulus raising, with optional support for a recently proposed technique (Bergerat et al, CHES 2025) based on the common mask assumption (CM packing). Our design keeps the accumulator update chain within a single thread block and fuses NTT/INTT, external products, and accumulator updates using a fixed execution template. Two compile-time parameters—WPP (warps per polynomial) and IPT (items per thread)—control multi-warp cooperation and per-thread register footprint, enabling consistent kernel structure across different parameter sets. On an NVIDIA RTX 4090, BOLT-FHE reaches 40,166 bootstrappings per second at 128-bit security, demonstrating high-throughput TFHE bootstrapping on a commodity GPU. Compared to the state-of-the-art GPU implementation VeloFHE (Shen et al, CHES 2025), BOLT-FHE achieves 1.85×–2.92× speedups with gadget decomposition. In particular, for modulus raising, BOLT-FHE improves by 2.38×–2.42× without CM packing, and by 3.17×–3.31× under the best packing configuration, reflecting the combined benefits of fused arithmetic, more regular memory access, and amortization enabled by CM packing. Overall, BOLT-FHE shows that a portable, fused-kernel organization with explicit on-chip budgeting can substantially improve TFHE bootstrapping throughput while remaining compatible with both noise management paths.
Note: Minor revision of the published TCHES version. This revision adds the open-source implementation and updates the performance results for the uint32-compatible parameter sets.
Metadata
- Available format(s)
-
PDF
- Category
- Implementation
- Publication info
- A minor revision of an IACR publication in TCHES 2026
- DOI
- 10.46586/tches.v2026.i3.198-223
- Keywords
- Fully Homomorphic EncryptionTFHEGPU Implementation
- Contact author(s)
-
yanrenchen @ mail ustc edu cn
zhengfangyu @ ucas ac cn
fanguang328 @ gmail com
djiankuo @ gmail com
wenxutang @ mail ustc edu cn
weekdayzt @ mail ustc edu cn
linjq @ ustc edu cn
jwjing @ ucas ac cn - History
- 2026-10-07: approved
- 2026-10-05: received
- See all versions
- Short URL
- https://ia.cr/2026/2356
- License
-
CC BY
BibTeX
@misc{cryptoeprint:2026/2356,
author = {Yanren Chen and Fangyu Zheng and Guang Fan and Jiankuo Dong and Wenxu Tang and Tian Zhou and Jingqiang Lin and Jiwu Jing},
title = {{BOLT}-{FHE}: An Efficient Unified Framework for {GPU}-based {TFHE} Bootstrapping via On-Chip Local Tiling Strategies},
howpublished = {Cryptology {ePrint} Archive, Paper 2026/2356},
year = {2026},
doi = {10.46586/tches.v2026.i3.198-223},
url = {https://eprint.iacr.org/2026/2356}
}