Timeline

Reddit sues Perplexity and scraping firms over AI data collection

Reddit alleged Perplexity used scraping firms to pull its posts indirectly from Google search results rather than pay for a licence, as OpenAI and Google had.

  • Courts & copyright
  • Notable

Reddit sued Perplexity and three data-scraping firms — SerpApi, Oxylabs and AWMProxy — in the US District Court for the Southern District of New York, alleging what it called “industrial-scale scraping” of Reddit posts and comments to build Perplexity’s AI-powered answer engine. Unlike OpenAI and Google, which had each struck paid licensing agreements for Reddit data, Reddit alleged Perplexity instead relied on the named scraping firms to pull Reddit content indirectly out of Google’s search results and resell or reuse it for AI purposes, sidestepping both a direct licence and Reddit’s own technical access controls.

The complaint alleged violations of the Digital Millennium Copyright Act’s anti-circumvention provisions, alongside claims of unfair competition and unjust enrichment, and sought damages and an injunction blocking further use of Reddit’s data. Perplexity rejected the allegations, saying it did not train models on Reddit content and instead provided summaries and citations of public discussions; SerpApi said it “strongly disagreed” and would contest the suit.

The case extended a pattern in which Reddit, having converted its user-generated archive into a licensing asset through deals with OpenAI and Google, treated companies that accessed the same data without paying as free-riders rather than simply competitors. It sat alongside a wider run of 2025 suits against Perplexity over scraping and copyright, including from Britannica and Merriam-Webster in September and, later that year, the Chicago Tribune and New York Times — part of a broader legal reckoning over whether search-style AI answer engines needed to license the content they summarised.