Home / Tech / Project Panama: Why an AI Company Anthropic Bought and Destroyed Millions of Books

Project Panama: Why an AI Company Anthropic Bought and Destroyed Millions of Books

Project Panama: Why an AI Company Bought and Destroyed Millions of Books

Artificial intelligence companies are competing in one of the most expensive technology races in history. While powerful chips and massive data centers often dominate headlines, another critical resource has quietly become the industry’s biggest strategic asset: high-quality books.

At the center of this debate is Project Panama, a confidential initiative by AI company Anthropic that reportedly purchased, scanned, and destroyed millions of physical books to train its Claude AI models. What began as an internal data acquisition project has now become one of the most controversial stories in AI development.

What Is Project Panama?

Project Panama was an internal Anthropic program launched in 2024 to build one of the world’s largest collections of high-quality digital books for AI training.

Internal planning documents reportedly described it as an effort to “destructively scan all the books in the world.” The project only became public in 2026 after court documents related to a copyright lawsuit were unsealed.

Unlike traditional digitization projects that preserve original books, Project Panama focused on speed and scale.

How the Project Worked

Anthropic reportedly invested tens of millions of dollars to purchase millions of second-hand books from bookstores, wholesalers, libraries, and bulk book suppliers.

The workflow was industrial:

  • Physical books were legally purchased in bulk.
  • Hydraulic cutting machines removed the bindings.
  • Individual pages were separated and scanned using high-speed commercial scanners.
  • The pages were converted into machine-readable digital files.
  • The remaining paper was recycled or discarded.

Once digitized, the original books no longer existed.

Vendor proposals reportedly estimated the ability to process 500,000 to 2 million books within six months, highlighting the unprecedented scale of the operation.

Why Did Anthropic Destroy Physical Books?

The answer lies in one word: training data.

Modern Large Language Models (LLMs) like Claude, ChatGPT, Gemini, and Grok learn from enormous collections of text.

Books offer several advantages over internet data:

  • Professionally edited writing
  • Reliable grammar and structure
  • Long-form reasoning
  • Scientific and historical accuracy
  • Rich vocabulary
  • Complex human arguments

Compared with noisy web pages, social media posts, and automatically generated content, books remain one of the highest-quality datasets available for AI.

Anthropic reportedly believed that access to millions of professionally written books would significantly improve Claude’s reasoning and writing abilities.

The Google Books Connection

To lead Project Panama, Anthropic hired Tom Turvey, a technology executive who previously played a major role in Google’s landmark Google Books digitization project.

Unlike Google Books—which largely relied on non-destructive scanning and returned borrowed books—Project Panama prioritized faster, lower-cost destructive scanning, sacrificing the physical copies after digitization.

Why Didn’t Anthropic Just Buy Ebooks?

Digital books usually come with licensing restrictions.

Physical books, however, fall under the first-sale doctrine in U.S. copyright law, allowing owners to resell, donate, modify—or even destroy—their legally purchased copies.

Anthropic used this legal distinction to create digital training copies from books it had purchased outright.

The Copyright Lawsuit

Project Panama became public during a high-profile copyright lawsuit brought by authors against Anthropic.

The case also revealed that, before launching Project Panama, Anthropic had allegedly obtained millions of unauthorized digital books from online shadow libraries such as LibGen and Books3—a far riskier legal approach.

Project Panama represented a shift toward legally purchased physical books instead of relying on pirated digital collections.

In a landmark decision, U.S. District Judge William Alsup ruled that scanning lawfully purchased books for internal AI training qualified as fair use because the process was considered transformative, even though the physical books were destroyed.

However, the court distinguished this from the earlier use of pirated material, which remained legally problematic. Anthropic later agreed to a reported $1.5 billion settlement related to copyright claims.

Why Project Panama Matters

Project Panama highlights a major shift in the AI industry.

The bottleneck is no longer just computing power—it is access to premium human-created knowledge.

As internet content becomes saturated with AI-generated material, books have become increasingly valuable because they contain:

  • Decades of expert knowledge
  • Editorial quality
  • Scientific research
  • Literature
  • Philosophy
  • Human creativity

This has sparked fierce competition among frontier AI companies to secure exclusive access to high-quality datasets.

Ethical Questions Raised

The project has divided researchers, authors, and publishers.

Critics argue that:

  • Destroying physical books reduces cultural preservation.
  • Authors deserve ongoing compensation when their work trains commercial AI systems.
  • Mass digitization without licensing could weaken the publishing industry.
  • Rare editions and annotated copies could be permanently lost.

Supporters counter that:

  • The books were legally purchased.
  • AI training is transformative rather than substitutive.
  • Digitization preserves knowledge in a searchable digital form.
  • Better AI systems can ultimately benefit education, science, and society.

The Bigger AI Data Arms Race

Project Panama is not an isolated case—it represents a broader trend across the AI industry.

Leading companies, including Anthropic, OpenAI, Google, Meta, xAI, and others, are increasingly competing for high-quality training data through licensing deals, public-domain archives, synthetic datasets, and legally acquired collections.

The race toward more capable AI is no longer driven solely by GPUs or larger models. Superior training data has become a strategic advantage. Models trained on richer, better-curated datasets are expected to produce stronger reasoning, fewer hallucinations, and better performance across coding, research, medicine, and scientific discovery.

Final Thoughts

Project Panama may become one of the defining moments in AI history. It exposed the extraordinary lengths technology companies are willing to go to obtain high-quality training data and highlighted the growing tension between innovation, copyright, and cultural preservation.

Whether remembered as a breakthrough in AI development or a cautionary tale about the cost of technological progress, Project Panama has already changed the conversation around how the next generation of AI models will be built—and who ultimately owns the knowledge that powers them.

Tagged: