
AI Companies Face Backlash Over Book Destruction for Training Data
AI companies seek physical books for training data, sparking ethical concerns over the destruction of literature and copyright implications.
The Rise of AI Companies Sourcing Training Data from Books
In recent years, artificial intelligence (AI) companies have increasingly turned their attention to a largely untapped resource for training data: physical books. This shift has led to significant ethical concerns regarding the destruction of these texts, as developers seek authentic human writing to train their models. Notably, ISBNdb briefly promoted a service aimed at providing large quantities of rare and specialized books for AI developers, underscoring a growing trend that many now liken to literary purges throughout history.
A Controversial Practice
ISBNdb's controversial sales pitch claimed to offer up to 1 million physical books tailored for AI training, particularly focusing on older, out-of-print titles. Their marketing promised confidentiality and strict nondisclosure agreements—elements that raised eyebrows among scholars and bibliophiles alike. When it was revealed that converting these physical books into digital format involved destructive scanning practices—essentially sacrificing the original works—public backlash ensued.
Despite the outcry, the demand for older books stems from a desire for untainted human literature, as the internet becomes saturated with machine-generated texts. AI developers aim to create comprehensive datasets that avoid the pitfalls of “model collapse,” a phenomenon wherein models trained on previous AI outputs begin to propagate inaccuracies.
Why Older Texts Matter to AI Developers
Large language models (LLMs) thrive on vast quantities of text, predominantly sourced from digital platforms. However, the rise of AI-generated content has made it increasingly difficult for developers to guarantee the authenticity of the datasets they collect. Older physical books, with their curation and peer-reviewed knowledge, provide an irreplaceable alternative.
ISBNdb highlighted this aspect in its marketing, noting that many potentially useful titles have never been fully digitized. Consequently, the urge to exploit these archival treasures has grown stronger as developers look for a clear avenue toward sourcing authentic human writing.
Historical Context of Book Digitization
The practice of digitizing written works is not new. Projects like Project Gutenberg, which began in 1971, have long sought to convert public-domain texts into accessible formats. This trend escalated notably with Google Books, which scanned libraries worldwide. However, these efforts generally prioritized preservation and access rather than exchanging physical copies for digital files.
In contrast, companies like Anthropic have taken more aggressive approaches. Documents show that Anthropic acquired millions of books, some of which were irreparably destroyed during scanning processes in a pursuit for vast sources of original human writing. As AI's quest for authentic content continues, it raises alarms about the potential loss of literary diversity.
The Legal Landscape and Copyright Concerns
The process of transforming physical books into digital files has instigated a complex legal landscape. In June 2025, a federal court ruled in favor of Anthropic, establishing that transforming purchased print books into digital formats constitutes fair use under specific conditions. However, it also drew distinctions between legitimately purchased materials and the vast quantities of pirated books that had previously been downloaded by the company.
This ruling underscored the precarious balance between intellectual property rights and the burgeoning demands of technology firms seeking to innovate rapidly. Authors and publishers are increasingly advocating for better protections regarding how their works are used in AI training. There is a growing call for negotiation in licensing agreements that ensure compensations for authors who find their works utilized in commercial AI products without proper attribution.
Public Reaction and Ethical Implications
As discussions around the destructive scanning of books have proliferated, the implications have resonated deeply with the public. Comparisons to infamous historical book burnings have come to the forefront, evoking anxiety over the loss of cultural heritage. These connections echo through literary communities, where the act of destroying books symbolizes censorship and the loss of intellectual diversity, as seen in Fahrenheit 451.
Conversely, some argue that repurposing unwanted or unsold books for digital preservation could serve a greater good, offering a lifeline for works that would otherwise be lost.
Conclusion: The Future of Books and AI
As AI companies continue to chase the bounty offered by the bookshelves while grappling with ethical dilemmas and legal frameworks, a question remains: how can developers balance their needs with the responsibility to preserve our literary heritage? The challenges they face will reflect in the future landscape of both AI technology and the world of published works, demanding innovative solutions that respect both literary and copyright principles while striving for technological advancement.
AI’s return to the bookshelf reflects a complex interplay of innovation and conservation. Not only do AI firms need to cultivate pathways for authentic human content, but they also bear the weight of responsibility toward the preservation of the written word.
Popular news
At least 11 miners are dead and dozens trapped following a coal mine explosion in Balochistan, Pakistan, raising safety concerns.
Subscribe to
our news
Get the most important updates and top stories in your inbox.





