Paper 2025/2251

An Efficient Private GPT Never Autoregressively Decodes

Zhengyi Li, Shanghai Jiao Tong University
Yue Guan, Shanghai Jiao Tong University
Kang Yang, State Key Laboratory of Cryptology
Yu Feng, Shanghai Jiao Tong University
Ning Liu, Shanghai Jiao Tong University
Yu Yu, Shanghai Jiao Tong University, Shanghai Qizhi Institute
Jingwen Leng, Shanghai Jiao Tong University, Shanghai Qizhi Institute
Minyi Guo, Shanghai Jiao Tong University, Shanghai Qizhi Institute
Abstract

The wide deployment of the generative pre-trained transformer (GPT) has raised privacy concerns for both clients and servers. While cryptographic primitives can be employed for secure GPT inference to protect the privacy of both parties, they introduce considerable performance overhead. To accelerate secure inference, this study proposes a public decoding and secure verification approach that utilizes public GPT models, motivated by the observation that securely decoding one and multiple tokens takes a similar latency. The client uses the public model to generate a set of tokens, which are then securely verified by the private model for acceptance. The efficiency of our approach depends on the acceptance ratio of tokens proposed by the public model, which we improve from two aspects: (1) a private sampling protocol optimized for cryptographic primitives and (2) model alignment using knowledge distillation. Our approach improves the efficiency of secure decoding while maintaining the same level of privacy and generation quality as standard secure decoding. Experiments demonstrate a $2.1\times \sim 6.0\times$ speedup compared to standard decoding across three pairs of public-private models and different network conditions.

Metadata
Available format(s)
PDF
Category
Cryptographic protocols
Publication info
Published elsewhere. ICML 2025
Keywords
secure multi-party computationhomomorphic encryptionTransformerLLMSecure Inferencespeculative decoding
Contact author(s)
hobbit @ sjtu edu cn
History
2025-12-18: revised
2025-12-15: received
See all versions
Short URL
https://ia.cr/2025/2251
License
Creative Commons Attribution
CC BY

BibTeX

@misc{cryptoeprint:2025/2251,
      author = {Zhengyi Li and Yue Guan and Kang Yang and Yu Feng and Ning Liu and Yu Yu and Jingwen Leng and Minyi Guo},
      title = {An Efficient Private {GPT} Never Autoregressively Decodes},
      howpublished = {Cryptology {ePrint} Archive, Paper 2025/2251},
      year = {2025},
      url = {https://eprint.iacr.org/2025/2251}
}
Note: In order to protect the privacy of readers, eprint.iacr.org does not use cookies or embedded third party content.