Skip to content
News

Anthropic trained AI on Dutch bestsellers using pirated library data

A major U.S. court settlement has exposed how artificial intelligence developers used copyrighted books by prominent Dutch writers to train their language models.

Essentially Amsterdam staff · Published July 23, 2026 at 1:52 p.m. CEST · 2 min read

ea-63d60e49

American artificial intelligence developer Anthropic utilized copyrighted books from well-known Dutch authors to train its language models without securing permission from the creators.

The training data was drawn from unauthorized digital repositories known as shadow libraries, which store large collections of pirated books and academic literature online.

Classic works included in training set

Among the affected works is the 2009 novel Het Diner by Herman Koch, alongside major classic titles by the late author Harry Mulisch, including De Aanslag and De Ontdekking van de Hemel.

The inclusion of these texts came to light through court documents connected to an ongoing class-action lawsuit filed against the tech company in the United States.

To resolve the legal claims, a federal judge in San Francisco approved a 1.5 billion dollar settlement requiring Anthropic to compensate affected authors.

Payout eligibility remains limited

Under the terms of the agreement, the company agreed to pay approximately 3,000 dollars for each book utilized during the training process.

However, many Dutch writers whose books were included in the database may not be eligible to collect a share of the settlement funds.

Dutch literary representatives note that international creators must generally be registered with the United States Copyright Office to receive compensation, a process usually tied to publishing an English translation in the American market.

Debate over artificial intelligence copyright

The settlement highlights broader disputes between technology companies and the creative sector over how digital models process published literature.

Tech firms have frequently argued that using existing texts to train algorithm models qualifies as fair use under current legal frameworks.

The issue remains closely watched by writers and publishers across the Netherlands as artificial intelligence tools continue to expand their reach.

Source: NL Times

More news