Paper 2024/1869

Black-box Collision Attacks on Widely Deployed Perceptual Hash Functions

Diane Leblanc-Albarel, KU Leuven
Bart Preneel, KU Leuven
Abstract

Perceptual hash functions have been designed to detect multimedia copyright violations and illegal content. To achieve their purpose, they have to maps inputs that are perceived as similar to close outputs. For several widely used designs the design strategy or even the design details are proprietary. Governments consider extending these functions to Client-Side Scanning (CSS) for end-to-end encrypted services, verifying content against illegal material before encryption. In 2021, Apple presented a detailed proposal for CSS based on the NeuralHash perceptual hash function. After strong criticism pointing out privacy and security concerns, Apple withdrew the proposal, but NeuralHash remains deployed on all devices, with its current purpose undisclosed. Brute-force collisions for NeuralHash (with a 96-bit result) require $2^{48}$ evaluations. Shortly after NeuralHash’s release, researchers showed it is easy to craft perceptually dissimilar collisions, enabling false incrimination in CSS by sending an innocent image with the same hash as illegal content. This work shows a more serious weakness: when inputs are restricted to a set of human faces, random collisions are highly likely to occur in input sets of size $2^{16}$. Unlike the targeted attack, our black-box attacks require no knowledge of the hash function’s design. We also show a high false negative rate (pictures that should share the same hash but do not). We show the generality of our approach by attacking PhotoDNA, Microsoft’s widely deployed 1152-bit perceptual hash. Small input sets yield near-collisions after only $2^{14.6}$ or $2^{17}$ operations, depending on the threshold. These results show that current perceptual hash designs are unsuitable for large-scale client scanning, producing high false positive and false negative rates, and highlight the need to reassess their security and feasibility, particularly for large-scale applications where privacy risks and false positives have serious consequences.

Metadata
Available format(s)
PDF
Category
Attacks and cryptanalysis
Publication info
Preprint.
Keywords
Perceptual HashingCollisionsClient-Side ScanningNeuralHashPhotoDNACSAM detection
Contact author(s)
diane leblanc-albarel @ kuleuven be
bart preneel @ esat kuleuven be
History
2025-08-26: last of 3 revisions
2024-11-15: received
See all versions
Short URL
https://ia.cr/2024/1869
License
Creative Commons Attribution-NonCommercial-NoDerivs
CC BY-NC-ND

BibTeX

@misc{cryptoeprint:2024/1869,
      author = {Diane Leblanc-Albarel and Bart Preneel},
      title = {Black-box Collision Attacks on Widely Deployed Perceptual Hash Functions},
      howpublished = {Cryptology {ePrint} Archive, Paper 2024/1869},
      year = {2024},
      url = {https://eprint.iacr.org/2024/1869}
}
Note: In order to protect the privacy of readers, eprint.iacr.org does not use cookies or embedded third party content.