US District Judge Araceli Martínez-Olguín granted final approval on July 20 to Anthropic's $1.5 billion settlement with a certified class of book authors and publishers, closing the largest copyright class action in US history and delivering the first major resolution among the wave of lawsuits over AI training data. The judge signed off over objections from a subset of class members who argued the deal undercompensates rightsholders, ruling the settlement was fair, reasonable and adequate given the risks of continued litigation.
The case, Bartz v. Anthropic, was filed in August 2024 by authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson, who accused Anthropic of training Claude on pirated copies of their books. The suit turned on an earlier ruling that split the underlying conduct in two: training on books Anthropic had legally purchased and destructively scanned was fair use, but downloading roughly 7 million books from pirate libraries like LibGen and PiLiMi was not -- and that piracy alone was enough to expose Anthropic to statutory damages plaintiffs' attorneys argued could have reached well into the tens of billions had the case gone to a jury on willfulness.
Roughly 500,000 individual titles ended up qualifying for the class after removing duplicates and non-eligible works, and eligible rightsholders are guaranteed a minimum of about $3,000 per title, split among co-authors where applicable. More than 91% of eligible claimants have already filed, according to Anthropic and the Authors Guild -- an unusually high participation rate suggesting most rightsholders saw the deal as a better outcome than years of appeals.
“Model companies that can show a clean-sourced or licensed corpus now have a concrete competitive argument to make to enterprise customers and acquirers alike.”
Anthropic is the first frontier lab to actually close one of these cases, and it isn't alone in facing them: The New York Times' suit against OpenAI and Microsoft over news content is still contested, Meta is defending a similar authors' suit, Getty Images has parallel claims against Stability AI in the US and UK, and Suno and Udio are fighting the major record labels over AI-generated music. Anthropic's settlement is now the first real market price for pirated-training-data liability, and every plaintiff's lawyer in those other cases just got a number to anchor on.
$1.5 billion sounds enormous until it's set against the balance sheet it's landing on: Anthropic reportedly filed confidentially for an IPO this month at a valuation north of $965 billion, with a revenue run rate reported near $47 billion. For a company at that scale, this settlement is a resolved liability, not an existential one -- which is exactly why clearing it out of a prospectus's risk factors matters more than the dollar figure does.
For founders and GPs, the settlement previews what training-data diligence will look like going forward: LPs and acquirers are going to start asking foundation-model companies directly where their pretraining corpus came from, whether any of it was scraped from shadow libraries, and what reserve exists for litigation. Model companies that can show a clean-sourced or licensed corpus now have a concrete competitive argument to make to enterprise customers and acquirers alike.
What to watch: whether the objecting authors appeal the final approval, given some argued $3,000 a title is too low against statutory maximums that can run to $150,000 per willful infringement; how the OpenAI, Meta and Getty cases price against this new comp; whether this affects Anthropic's IPO timeline or prospectus disclosures; and whether Congress or the Copyright Office treat a $1.5 billion price on piracy as a mandate to legislate a compulsory licensing regime for AI training data.