On August 28, 2026, Sony Music Publishing, Warner Chappell Music, and more than thirty other music publishers filed a sweeping federal copyright lawsuit against Anthropic — the AI company behind the chatbot Claude — alleging that the company built its models on a foundation of pirated songs. The case is Sony Music Publishing (US) LLC v. Anthropic PBC, No. 5:26-cv-09217, filed in the Northern District of California and assigned to Judge Edward M. Chen, under the Copyright Act (17 U.S.C. § 101 et seq.). With more than 120 articles across 115 outlets tracked in its first days, it is one of the most heavily covered new IP filings on the IP Feed.
It is also the clearest test yet of a question the entire generative-AI industry is circling: when a model is trained on copyrighted work, is the copying that made it possible fair use — or infringement on a catalog-sized scale?
What the lawsuit says
The publishers allege that Anthropic "unlawfully acquired and copied" thousands of copyrighted musical compositions — lyrics and sheet music — by "torrenting, scraping, and downloading" them, and then used those copies to develop and train its Claude models. The 48-page complaint frames this not as an incidental byproduct of web-scale training but as a deliberate choice to build on infringing copies rather than license them.
The named works read like a radio programmer's greatest-hits list: the complaint points to songs including "Ain't No Mountain High Enough," "All I Want for Christmas Is You," "Eye of the Tiger," "Livin' On a Prayer," "September," and "Hallelujah." Beyond how the works were obtained, the publishers allege that Claude can reproduce copyrighted lyrics close to verbatim when prompted — an output-side harm distinct from the training-side copying.
The suit does not stop at the corporate defendant. It names Anthropic's chief executive, Dario Amodei, and co-founder Benjamin Mann as individual defendants. And it comes with real financial teeth: the plaintiffs seek statutory damages of up to $150,000 for each work willfully infringed, plus up to $25,000 for each alleged removal of copyright management information — the metadata and attribution that identifies who owns a song. Across thousands of compositions those numbers compound into the "multi-billion-dollar" figure the coverage has fixated on; the complaint calls it "one of the largest and most blatant ongoing thefts of intellectual property in history."
None of these allegations has been tested. As of this writing Anthropic has not answered the complaint, no motion has been briefed, and nothing has been ruled on. What follows is the legal terrain the filing sits on — not a prediction of how it comes out.
The hard part: training, sourcing, and output
Cases like this one are easy to describe in a headline and hard to resolve in a courtroom, because they actually contain three different questions that courts treat very differently.
- The training question. Does copying a work in order to train a model — to extract statistical patterns rather than to republish the work — qualify as transformative fair use? This is the argument AI developers have leaned on, and in the earlier wave of author litigation against Anthropic a federal court found that training itself could be transformative. That holding is the industry's strongest card.
- The sourcing question. Where did the copies come from? That same author litigation drew a sharp line: training might be fair use, but downloading pirated copies to assemble a permanent library was a separate act that fair use did not excuse. The publishers' emphasis on "torrenting" is not rhetorical color — it is aimed squarely at that line, because how copies are acquired can be infringing regardless of what they are later used for.
- The output question. If Claude can be prompted to return protected lyrics nearly verbatim, that looks less like learning from a work and more like reproducing it — the fact pattern fair use is least forgiving of.
The publishers have structured their complaint to win on any one of these, not all three. Anthropic, for its part, will almost certainly argue transformative use, contest whether verbatim reproduction actually happens at scale in ordinary use, and dispute the provenance allegations. The distance between "a model learned from our songs" and "a company copied our catalog without a license" is exactly where this case will be fought.
Why it matters beyond one lawsuit
Music is a uniquely unforgiving arena for this fight. A song's lyrics and its composition are separately owned and heavily registered; publishers know precisely what they hold, and licensing catalogs to third parties is not a hypothetical — it is the business. That makes "you should have licensed this" a far more concrete argument than it is for loosely tracked web text, and it makes statutory damages, which do not require proving actual harm, genuinely dangerous at catalog scale.
For anyone who creates or owns rights, the case is a reminder that provenance is the whole ballgame. The exposure here turns less on whether AI training is inherently lawful and more on a paper trail: what was copied, where it came from, whether it was licensed, and whether the identifying information attached to each work was stripped along the way. The value of a creative work is inseparable from a clean, provable record of who owns it and how it was used — and disputes like this are what that record is for.
We'll track the docket as it develops — the first real signal will come when Anthropic responds and the shape of its defense (transformative use, disputed sourcing, or both) comes into view. Follow the filings, counsel, and coverage on the case page.
This post is editorial commentary on public court filings and news coverage, not legal advice. The allegations described are unproven, and the defendants have not yet responded in court.