Paper 2026/2356

BOLT-FHE: An Efficient Unified Framework for GPU-based TFHE Bootstrapping via On-Chip Local Tiling Strategies

Yanren Chen, School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China
Fangyu Zheng, School of Cryptology, University of Chinese Academy of Sciences, Beijing, China
Guang Fan, School of Cryptology, University of Chinese Academy of Sciences, Beijing, China
Jiankuo Dong, School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China
Wenxu Tang, School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China
Tian Zhou, School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China
Jingqiang Lin, School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China
Jiwu Jing, School of Cryptology, University of Chinese Academy of Sciences, Beijing, China
Abstract

Bootstrapping is the main performance bottleneck in bitwise Fully Homomorphic Encryption (FHE), and practical acceleration requires careful orchestration of the blind rotation and external product chain under GPU resource constraints. This paper presents BOLT-FHE, a GPU bootstrapping framework that emphasizes block-local execution, on-chip tiling, and a unified MegaKernel supporting both gadget decomposition and modulus raising, with optional support for a recently proposed technique (Bergerat et al, CHES 2025) based on the common mask assumption (CM packing). Our design keeps the accumulator update chain within a single thread block and fuses NTT/INTT, external products, and accumulator updates using a fixed execution template. Two compile-time parameters—WPP (warps per polynomial) and IPT (items per thread)—control multi-warp cooperation and per-thread register footprint, enabling consistent kernel structure across different parameter sets. On an NVIDIA RTX 4090, BOLT-FHE reaches 40,166 bootstrappings per second at 128-bit security, demonstrating high-throughput TFHE bootstrapping on a commodity GPU. Compared to the state-of-the-art GPU implementation VeloFHE (Shen et al, CHES 2025), BOLT-FHE achieves 1.85×–2.92× speedups with gadget decomposition. In particular, for modulus raising, BOLT-FHE improves by 2.38×–2.42× without CM packing, and by 3.17×–3.31× under the best packing configuration, reflecting the combined benefits of fused arithmetic, more regular memory access, and amortization enabled by CM packing. Overall, BOLT-FHE shows that a portable, fused-kernel organization with explicit on-chip budgeting can substantially improve TFHE bootstrapping throughput while remaining compatible with both noise management paths.

Note: Minor revision of the published TCHES version. This revision adds the open-source implementation and updates the performance results for the uint32-compatible parameter sets.

Metadata
Available format(s)
PDF
Category
Implementation
Publication info
A minor revision of an IACR publication in TCHES 2026
DOI
10.46586/tches.v2026.i3.198-223
Keywords
Fully Homomorphic EncryptionTFHEGPU Implementation
Contact author(s)
yanrenchen @ mail ustc edu cn
zhengfangyu @ ucas ac cn
fanguang328 @ gmail com
djiankuo @ gmail com
wenxutang @ mail ustc edu cn
weekdayzt @ mail ustc edu cn
linjq @ ustc edu cn
jwjing @ ucas ac cn
History
2026-10-07: approved
2026-10-05: received
See all versions
Short URL
https://ia.cr/2026/2356
License
Creative Commons Attribution
CC BY

BibTeX

@misc{cryptoeprint:2026/2356,
      author = {Yanren Chen and Fangyu Zheng and Guang Fan and Jiankuo Dong and Wenxu Tang and Tian Zhou and Jingqiang Lin and Jiwu Jing},
      title = {{BOLT}-{FHE}: An Efficient Unified Framework for {GPU}-based {TFHE} Bootstrapping via On-Chip Local Tiling Strategies},
      howpublished = {Cryptology {ePrint} Archive, Paper 2026/2356},
      year = {2026},
      doi = {10.46586/tches.v2026.i3.198-223},
      url = {https://eprint.iacr.org/2026/2356}
}
Note: In order to protect the privacy of readers, eprint.iacr.org does not use cookies or embedded third party content.