Timeline

Stanford researchers find CSAM links in LAION-5B, dataset pulled

Scanning roughly 32 million data points, researchers flagged 3,226 suspected matches and externally validated 1,008 using hash-matching services before LAION withdrew the dataset.

  • Open weights & ecosystem
  • Security & misuse
  • Notable

Researchers at the Stanford Internet Observatory, led by David Thiel, published an investigation into LAION-5B, the roughly 5-billion-image dataset used to train Stable Diffusion and other image generators, and reported finding child sexual abuse material within it. Scanning a subset of around 32 million data points using perceptual and cryptographic hash-matching against databases maintained by child-safety organisations, the investigation flagged 3,226 suspected instances and externally validated 1,008 of them through the hash-matching services PhotoDNA and others, with the researchers describing that count as a likely undercount given the scanning method’s limitations.

LAION, the German nonprofit that maintains the dataset, said it removed LAION-5B and the related LAION-400M dataset from circulation “out of an abundance of caution,” pending safety review before any republication. The organisation said it had existing filters intended to screen out illegal content but had not been able to prevent all such material from entering a dataset built by scraping billions of images and captions from the public web.

The finding raised a distinct concern from most AI dataset controversies to that point, which had centred on consent and copyright rather than dataset legality: the Stanford team noted that models trained on LAION-5B, including widely deployed versions of Stable Diffusion, may have learned from and could reproduce material derived from the abuse images, and that anyone who had downloaded the full dataset likely possessed CSAM without realising it, absent specific precautions. The report did not claim that image generators trained on the dataset had been shown to reliably reproduce identifiable victims, but flagged reinforcement of harm to those victims as a distinct risk from the dataset’s mere existence.

The episode prompted LAION to release a filtered successor dataset, Re-LAION-5B, the following year, and added dataset provenance and safety screening to the list of scrutiny generative-image companies faced alongside the copyright disputes already underway.