Timeline

The Authors Guild sues OpenAI

Seventeen novelists including Grisham and Martin filed a class action in the Southern District of New York alleging systematic copying of pirated books to train GPT.

  • Courts & copyright
  • Notable

The Authors Guild and seventeen novelists — including John Grisham, George R.R. Martin, Jodi Picoult, Jonathan Franzen, David Baldacci and Scott Turow — filed a proposed class action against OpenAI in the US District Court for the Southern District of New York, alleging that GPT-3.5 and GPT-4 had been trained on their copyrighted novels without permission or payment.

The complaint alleged what it called “systematic theft on a mass scale,” claiming books had been copied wholesale from pirate ebook repositories into OpenAI’s training data. It sought to represent a class of fiction writers whose work had been used this way, and argued the resulting models could produce derivative text that mimicked named authors’ styles, characters and voices — a capability the plaintiffs said threatened writers’ livelihoods directly, since a chatbot able to write “in the style of” a living novelist competes with that novelist’s own future output.

The suit joined a wave of author-led litigation that had begun that summer, but its list of bestselling, high-profile plaintiffs and the Authors Guild’s institutional backing gave it wider media attention than earlier individual filings. The case was later amended to add Microsoft as a defendant, and it became one of several parallel author actions against OpenAI working through the same court, alongside The New York Times’ suit filed that December.

None of these cases had resolved the underlying legal question — whether training a language model on copyrighted text is fair use — by the time they were filed; that determination fell to courts over the following years, with outcomes that varied by defendant and by the specific facts of how training data had been acquired. The Authors Guild suit was significant less for any immediate ruling than for establishing, within OpenAI’s first year of mass consumer adoption, that book publishing’s most visible authors and their trade association intended to litigate training-data use as copyright infringement rather than treat it as an unresolved grey area.