Organisation
LAION
A German non-profit that builds and releases large open image-text datasets used to train open generative-image models.
LAION is a German non-profit that assembles and releases large open datasets for training AI, most notably LAION-5B, a set of nearly six billion image-and-caption pairs scraped from the public web. That dataset became the training corpus for Stable Diffusion and many other open image-generation models, making it, for several years, a foundational piece of shared infrastructure in the open-model ecosystem. Its scale and lack of curation also made it a liability: in late 2023 Stanford researchers reported finding links to child sexual abuse material within it, prompting LAION to withdraw the dataset and later release a filtered successor, Re-LAION-5B. The episode added dataset provenance and safety screening to the scrutiny generative-image systems already faced over copyright.
- Category
- Open source & ecosystem
- Founded
- 2021
- HQ
- Hamburg, DE
- Key people
- Christoph Schuhmann, Jenia Jitsev
Featured in threads
Tracks
- Open weights & ecosystem 4
- Security & misuse 2
- Models & capabilities 1
LAION releases Re-LAION-5B with CSAM links removed
The cleaned dataset shrank from 5.8 to 5.5 billion pairs after matching hashes from the Internet Watch Foundation and the Canadian Centre for Child Protection, eight months after Stanford's report.
Open weights & ecosystem · Security & misuse
Stanford researchers find CSAM links in LAION-5B, dataset pulled
Scanning roughly 32 million data points, researchers flagged 3,226 suspected matches and externally validated 1,008 using hash-matching services before LAION withdrew the dataset.
Open weights & ecosystem · Security & misuse
Stable Diffusion is released to the public
Weights published under a permissive licence and runnable on a consumer graphics card, with community fine-tunes and graphical front-ends appearing within weeks.
Open weights & ecosystem · Models & capabilities
LAION releases LAION-5B image-text dataset
LAION released LAION-5B, then the largest freely available image-text dataset (5.85bn pairs), which became the training corpus for Stable Diffusion and other open image models.
Open weights & ecosystem
Also mentioned in 1 entry
Referenced in passing — LAION isn't the main subject of these.