Paper 2026/1578

OpenLLM: Modular and Scalable zkSNARKs for Verifiable LLM Inference

Yunbo Yang, The State Key Laboratory of Blockchain and Data Security, Zhejiang University, Hangzhou High-Tech Zone (Bin jiang) Institute of Blockchain and Data Security
Yupeng Ren, State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, School of Cyber Security, University of Chinese Academy of Sciences
Changtong Xu, Ant Group
Rui Zhang, State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering Chinese Academy of Sciences
Xuanming Liu, Zhejiang University
Jin Tan, Ant Group
Tao Wei, Ant Group
Bingsheng Zhang, Zhejiang University
Kui Ren, Zhejiang University
Abstract

Large language model (LLM) is increasingly deployed as a remote service, where users rely on third-party servers to perform computation. However, such settings introduce critical integrity concerns, as an untrusted server may deviate from the prescribed computation, skip expensive operations, or return incorrect results, while users lack practical approaches to verify execution correctness. Ensuring the correctness of LLM inference under untrusted execution remains a fundamental challenge. Zero-knowledge proofs (ZKPs) provide a principled approach verifying computation correctness, but applying them to LLM inference remains challenging. Modern LLMs involve a large number of non-linear operations and require modeling real-valued computation in finite fields, introducing substantial computational and memory overhead and potential loss of numerical precision. Moreover, the large scale of LLMs makes end-to-end verification difficult to scale, limiting the practicality of existing approaches. This paper presents OpenLLM, an efficient and modular system for verifiable LLM inference. Our key idea is to decompose large-scale LLM inference into a set of reusable atomic operators, each equipped with efficient ZKP protocols, enabling scalable verification at the operator level. Based on this abstraction, we design succinct non-interactive zero-knowledge proof constructions for representative non-linear functions, which can be composed into end-to-end inference pipelines independent of model architectures. We further evaluate OpenLLM across operator-level performance, end-to-end inference, layer-wise scaling, larger models, and approximation accuracy. The results show that OpenLLM achieves smaller proof sizes, lower verification cost, and improved numerical fidelity while scaling from individual operators to full-model inference. Compared with state-of-the-art interactive protocols, OpenLLM eliminates communication overhead through a fully non-interactive design while maintaining competitive efficiency. Building on this operator-level efficiency, it further enables a scalable and modular framework for end-to-end verifiable LLM inference, outperforming prior end-to-end approaches.

Metadata
Available format(s)
PDF
Category
Applications
Publication info
Preprint.
Keywords
Applied CryptographyMachine LearningAILarge Language ModelsZero-knowledge Proof
Contact author(s)
yyb9882 @ gmail com
renyupeng @ iie ac cn
xuchangtong xct @ antgroup com
zhangrui @ iie ac cn
hinsliu @ zju edu cn
tanjin tj @ antgroup com
lenx wei @ antgroup com
bingsheng @ zju edu cn
kuiren @ zju edu cn
History
2026-08-03: approved
2026-08-02: received
See all versions
Short URL
https://ia.cr/2026/1578
License
Creative Commons Attribution
CC BY

BibTeX

@misc{cryptoeprint:2026/1578,
      author = {Yunbo Yang and Yupeng Ren and Changtong Xu and Rui Zhang and Xuanming Liu and Jin Tan and Tao Wei and Bingsheng Zhang and Kui Ren},
      title = {{OpenLLM}: Modular and Scalable {zkSNARKs} for Verifiable {LLM}  Inference},
      howpublished = {Cryptology {ePrint} Archive, Paper 2026/1578},
      year = {2026},
      url = {https://eprint.iacr.org/2026/1578}
}
Note: In order to protect the privacy of readers, eprint.iacr.org does not use cookies or embedded third party content.