Paper 2026/2299

Can AI Oversight Be Zero Knowledge?

Alessandro Chiesa, École Polytechnique Fédérale de Lausanne
Ziyi Guan, Massachusetts Institute of Technology
Burcu Yıldız, École Polytechnique Fédérale de Lausanne
Abstract

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

Metadata
Available format(s)
PDF
Category
Foundations
Publication info
Preprint.
Keywords
zero knowledgerelativized argumentsdebateoracle-aided computationscalable oversight
Contact author(s)
alessandro chiesa @ epfl ch
ziyiguan @ mit edu
burcu yildiz @ epfl ch
History
2026-10-04: approved
2026-10-01: received
See all versions
Short URL
https://ia.cr/2026/2299
License
Creative Commons Attribution
CC BY

BibTeX

@misc{cryptoeprint:2026/2299,
      author = {Alessandro Chiesa and Ziyi Guan and Burcu Yıldız},
      title = {Can {AI} Oversight Be Zero Knowledge?},
      howpublished = {Cryptology {ePrint} Archive, Paper 2026/2299},
      year = {2026},
      url = {https://eprint.iacr.org/2026/2299}
}
Note: In order to protect the privacy of readers, eprint.iacr.org does not use cookies or embedded third party content.