OpenAI desperate to avoid explaining why it deleted pirated book datasets
1 December 2025 at 17:16
OpenAI may soon be forced to explain why it deleted a pair of controversial datasets composed of pirated books, and the stakes could not be higher.
At the heart of a class-action lawsuit from authors alleging that ChatGPT was illegally trained on their works, OpenAIβs decision to delete the datasets could end up being a deciding factor that gives the authors the win.
Itβs undisputed that OpenAI deleted the datasets, known as βBooks 1β and βBooks 2,β prior to ChatGPTβs release in 2022. Created by former OpenAI employees in 2021, the datasets were built by scraping the open web and seizing the bulk of its data from a shadow library called Library Genesis (LibGen).


Β© wenmei Zhou | DigitalVision Vectors