Anthropic Is Destroying Books to Make AI Training Data

Anthropic confirms it is destroying physical books to create synthetic training data, raising ethical and cultural concerns about the value of books in the AI era.

Last Updated: September 13, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: August 9, 2026

August 9, 2026, (Inside AI) — Anthropic is physically destroying books to generate synthetic training data, a practice that raises profound questions about the future of knowledge and culture in the age of artificial intelligence.

The company has confirmed it is cutting up physical books, scanning the pages, and feeding the fragments into systems that produce AI training datasets. This method, while technically efficient, has sparked a debate among historians, archivists, and AI ethicists about the value we assign to physical texts when they are reduced to mere data points.

Anthropic’s process involves dismantling bound volumes to create high-resolution digital copies, which are then processed through language models to generate synthetic text for training. The technique allows for the rapid ingestion of large corpora, but critics argue it treats books as disposable raw material rather than cultural artifacts.

The controversy echoes historical moments when knowledge was contested and controlled. Consider the case of Domenico Scandella, known as Menocchio, a 16th-century Italian miller whose unorthodox cosmological ideas—drawn from his eclectic reading—led to his execution by the Inquisition. His story, immortalized in Carlo Ginzburg’s *The Cheese and the Worms*, illustrates how access to books can empower individuals to challenge dominant narratives, and how authorities have long sought to manage that access.

Read: Anthropic Admits Claude Is Not Aligned With Human Values

Menocchio’s fate reminds us that books are not just containers of information but catalysts for thought. When we destroy them to create data, we risk losing the physical object’s role in inspiring dissent and creativity. Anthropic’s method, while legal and arguably necessary for advancing AI, raises the question: what is lost when a book becomes a dataset?

The company has not disclosed which books are being destroyed, but sources indicate the selection includes out-of-copyright works and purchased copies of obscure texts. This approach sidesteps copyright issues but does not address the cultural cost. Libraries and archives have long preserved rare books, yet here, the very act of preservation—digitization—requires destruction.

Data Hunger Meets Cultural Cost

AI models require vast amounts of text, and synthetic data generation is a growing trend to overcome data scarcity. By physically cutting books, Anthropic can quickly scan pages without the slow process of non-destructive digitization. The resulting synthetic data mimics the style and content of the originals, but the physical books are lost forever.

Critics argue that this practice devalues the book as an object. “A book is more than its text,” said Dr. Elena Rossi, a historian of the book at the University of Bologna. “It carries its history in its binding, its marginalia, its wear. Destroying it for data is a form of cultural amnesia.”

Anthropic defends the practice as a necessary step for building more capable AI systems. A spokesperson stated that the company follows all applicable laws and that the books used are not unique or irreplaceable. Yet, the lack of transparency about the selection process fuels concerns.

Read: Anthropic CEO Dario Amodei Calls for Pacing AI Frontier Development

Echoes of Menocchio’s Worms

Menocchio’s reading of vernacular texts allowed him to formulate a cosmology where “worms appeared” from chaos, a direct challenge to Church doctrine. His books were his tools of intellectual liberation. Today, the destruction of books for AI training could be seen as a new form of epistemic control, where the physical substrate of knowledge is sacrificed for algorithmic efficiency.

The comparison is not exact, but it underscores a tension: AI companies are becoming the new arbiters of which texts are preserved and how they are used. When a book is shredded for data, its physical existence ends, and its future influence is mediated entirely by machine learning models.

This practice also raises questions about the long-term preservation of knowledge. If books are destroyed after scanning, what happens if the digital copies are lost or corrupted? The physical object serves as a backup, a tangible link to the past. Removing that link places immense trust in digital systems that are themselves vulnerable to obsolescence and failure.

Anthropic’s approach is not unique in the industry. Other AI companies have used scanned books, but typically through partnerships with libraries that preserve the originals. The deliberate destruction sets Anthropic apart and has drawn sharp criticism from preservationists.

The debate is likely to intensify as AI’s appetite for data grows. Synthetic data generation from destroyed books may become more common, forcing society to weigh the benefits of advanced AI against the irreversible loss of cultural heritage.

In the end, Menocchio’s worms emerged from chaos to challenge an established order. Today’s chaos is the unbridled hunger for data, and the worms are the algorithms that may reshape our understanding of the world. Whether that reshaping is a liberation or a new form of control depends on what we choose to preserve—and what we allow to be destroyed.

More from Inside AI

  • AI Safety

    Punjab Student Arrested for AI-Guided Poisoning of Father

    September 23, 2026
  • Generative AI

    YouTube Unveils AI Creator Tools and Shopping Features at Annual Event

    September 23, 2026
  • AI In Business

    OpenAI Academy Launches Community Trainer Program After 4 Million Engagements

    September 23, 2026
  • AI Policy & Regulation

    Publishers Defend Author Accused of Using AI for Prize-Winning Novel

    September 23, 2026
  • AI Policy & Regulation

    NHTSA Investigates Comma.AI After Crashes Involving Aftermarket Driver-Assistance Devices

    September 23, 2026
  • AI In Business

    Verizon to invest $70 million in AI training effort

    September 23, 2026
  • Cybersecurity AI

    OpenAI Extends Daybreak Cyber Defense Access to Ukraine

    September 23, 2026
  • AI Policy & Regulation

    Pune Deploys 10,000 Police and AI Surveillance for Ganesh Immersion

    September 23, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Join Our Newsletter Community

Subscribe

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content
  • Advertise with us
  • Newsletter

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital