The Daily
Menu
Tech

Federal Court Approves Landmark $1.5 Billion Anthropic Copyright Settlement

A federal judge has finalized a $1.5 billion settlement between Anthropic and a group of authors and publishers over copyright infringement. While the payout resolves the immediate dispute, the broader legal question of whether AI training constitutes fair use remains unresolved.

By Ada

A federal judge has granted final approval to a $1.5 billion settlement between AI lab Anthropic and a collective of authors and publishers, marking a significant resolution in a high-profile copyright infringement case. According to TechCrunch, the agreement concludes a class-action lawsuit that centered on the methods Anthropic employed to build its training datasets for its AI models.

Claude AI interface | Source: TechCrunch
Claude AI interface | Source: TechCrunch

The legal conflict originated from the company's use of two distinct sources for its training library: books that were purchased and scanned, and materials obtained from unauthorized sources such as Library Genesis and Pirate Library Mirror. While the court ruled that the act of training an AI model on copyrighted text qualifies as fair use—a decision that provides a degree of clarity for the broader AI industry—the judge determined that the acquisition of books from pirate sites was legally impermissible. Anthropic opted to settle the case to avoid a jury trial regarding the damages associated with these specific acquisition methods.

Under the terms of the settlement, approximately $3,000 will be distributed per work for an estimated 500,000 copyrighted items. Despite the magnitude of the payout, which is considered the largest in U.S. copyright history, the resolution does not establish a binding legal precedent. Because the ruling originated in a district court and the case was settled before reaching an appellate level, other courts remain free to interpret the legality of AI training practices differently in future litigation.

The landscape of AI copyright remains highly volatile as other major technology firms face similar scrutiny. Industry leaders including Google, Meta, Midjourney, and OpenAI are currently navigating their own legal challenges regarding the use of copyrighted data in model development. Most recently, a coalition of publishers and authors, including Hachette and Elsevier, initiated a new class-action lawsuit against Google, alleging that the company utilized protected works to train its Gemini platform. As these cases proceed through the court system, the industry continues to operate in a state of legal uncertainty, with the ultimate definition of fair use in the age of generative AI still waiting to be codified by higher courts.