SAN FRANCISCO — In what legal experts are already calling a watershed moment for artificial intelligence and intellectual property law, two of the world’s largest music publishers filed a sweeping copyright infringement lawsuit against AI powerhouse Anthropic and two of its top executives.
Sony Music Publishing and Warner Chappell Music launched the legal offensive on August 28, 2026, in the U.S. District Court for the Northern District of California. Filed collectively on behalf of a broad coalition of music publishers, the lawsuit alleges that Anthropic systematically copied tens of thousands of copyrighted musical compositions without authorization to train its flagship Claude AI models.
The complaint does not merely target the corporate entity. It also names Anthropic Chief Executive Officer Dario Amodei and co-founder Benjamin Mann as individual defendants, alleging they played direct roles in orchestrating the unauthorized acquisition of copyrighted text and sheet music.
Describing the alleged conduct as “one of the largest and most blatant ongoing thefts of intellectual property in history,” the plaintiffs are demanding a jury trial. The high-profile catalog of allegedly infringed works spans generations of popular culture, including classics like “Ain’t No Mountain High Enough,” “All I Want for Christmas is You,” “Eye of the Tiger,” “Livin’ On a Prayer,” “September,” and “Hallelujah,” alongside contemporary hits such as Taylor Swift’s “Paper Rings.”
Executive Overview: A High-Stakes Collision Between Music and Machine Learning
The legal action marks a profound escalation in the ongoing war between generative AI developers and creative industries. While visual artists, authors, and news publishers have waged protracted legal battles against major tech platforms over data scraping and model training, the music publishing sector has increasingly drawn a hard line in the digital sand.
At the heart of the lawsuit is a familiar, contentious question: Does the ingestion of copyrighted works to train commercial artificial intelligence models constitute fair use, or is it systemic, large-scale copyright infringement?
Unlike previous lawsuits that focused primarily on web-scraped text or unstructured datasets, the complaint against Anthropic introduces a darker narrative. Drawing heavily on unsealed internal records from a prior class-action lawsuit, the publishers allege that Anthropic’s data acquisition strategies crossed the line from aggressive web scraping into deliberate, industrial-scale piracy.
The plaintiffs argue that Anthropic knowingly utilized illicitly obtained repositories—including notorious shadow libraries and peer-to-peer torrent networks—to build the foundational datasets that power Claude. Furthermore, the lawsuit contends that Anthropic’s engineering practices deliberately scrubbed metadata and copyright notices to conceal the origins of the training data, while the resulting AI models continue to output memorized lyrics and market-diluting derivatives.
Detailed Chronology and Technical Allegations
The complaint outlines a four-count legal strategy, detailing how Anthropic allegedly acquired, processed, and utilized copyrighted musical compositions across multiple vectors.
The Four Core Counts
- Direct Copyright Infringement (Torrenting): Filed against Anthropic, Amodei, and Mann for unauthorized peer-to-peer downloading and distribution.
- Contributory Infringement: Charged personally against Amodei and Mann for directing, approving, and overseeing the torrenting operations.
- Direct Infringement (Scraping and Scanning): Brought against Anthropic alone for web scraping licensed lyric sites, executing physical book scanning operations, and utilizing third-party datasets.
- Copyright Management Information (CMI) Violations: Alleging that Anthropic intentionally removed or altered tracking data, such as song titles, songwriter names, and copyright notices, in violation of federal law.
From "Sketchy AF" Torrents to Physical Book Scanning
According to the legal filing, the foundation of Anthropic’s data collection involved deep incursions into digital shadow libraries. The complaint alleges that in June 2021, co-founder Benjamin Mann used the BitTorrent protocol to download at least five million pirated books from Library Genesis (LibGen). Shortly after, in July 2022, Anthropic employees allegedly pulled at least two million additional volumes from a successor repository known as Pirate Library Mirror (PiLiMi).
The publishers state that these digital hauls included hundreds of songbooks and sheet-music collections containing their proprietary compositions. Because the BitTorrent protocol inherently requires users to upload fragments of files to others while downloading, the lawsuit argues that every single torrent operation constituted an independent violation of the publishers’ exclusive right of distribution.
Beyond peer-to-peer networks, the complaint details a multi-pronged acquisition strategy:
- Lyric Scraping: Anthropic allegedly scraped lyrics directly from licensed websites like MusixMatch and LyricFind, flagrantly violating those platforms’ terms of service.
- Destructive Physical Scanning: The company reportedly operated an internal scanning initiative that digitized millions of second-hand physical books before destroying the source materials.
- Aggregated Datasets: Claude’s training pipelines allegedly drew from contentious third-party data repositories, including Common Crawl, The Pile, and Books3.
The Smoking Gun: Reliance on Bartz v. Anthropic
Much of the evidentiary weight in the new music publishers’ complaint borrows from Bartz v. Anthropic, a landmark authors’ class action litigated in the same federal district. In Bartz, the presiding court condemned Anthropic’s torrenting activities as “straightforward piracy but at massive scale,” a case that Anthropic ultimately settled in September 2025 for a staggering $1.5 billion.
The music publishers’ complaint weaponizes unsealed internal communications from the Bartz proceedings to demonstrate willful infringement. Among the cited internal exhibits:

- Co-founder Benjamin Mann characterized Library Genesis in internal chats as “sketchy AF.”
- An Anthropic archive team explicitly recognized the platform as a “blatant violation of copyright.”
- A 2024 internal planning document outlining the physical book-scanning initiative explicitly warned: “We don’t want it to be known that we are working on this.”
“Dr. Amodei and Mr. Mann are personally liable for their respective roles in this illegal torrenting of pirated copies of Music Publishers’ works from LibGen and PiLiMi,” the complaint asserts, piercing the corporate veil to hold leadership accountable.
Training Pipelines, Outputs, and the Limits of Guardrails
The lawsuit provides a granular look at how ingested data transforms into commercial AI capabilities. When Anthropic builds its training corpus, engineers allegedly run the text through automated "cleaning" and extraction tools. These utilities strip out copyright notices, publisher names, and author credits while retaining the underlying expressive content—a process the publishers characterize as deliberate concealment.
Once trained, Claude models allegedly memorize these protected lyrics. When prompted by users, the AI can reproduce the lyrics verbatim or near-verbatim. Furthermore, the model is capable of generating derivative works "in the style of" represented songwriters.
While Anthropic implemented text-generation guardrails following earlier legal challenges to prevent Claude from spitting out copyrighted lyrics, the publishers argue these measures are superficial. According to the filing, users can easily bypass the safeguards by simply re-prompting the model with slight variations.
From an economic perspective, the plaintiffs argue that Claude’s ability to generate custom lyrics creates direct market substitutes that compete with the publishers’ catalogs. This substitution effect, they contend, dilutes the broader streaming royalty pools from which both publishers and songwriters derive their livelihoods.
Supporting Context, Metrics, and Damages Demanded
The financial stakes of the lawsuit reflect the massive valuations and capital expenditures defining the generative AI sector.
The plaintiffs are pursuing maximum statutory remedies under U.S. copyright law:
- Statutory Damages: Up to $150,000 per infringed work for instances where infringement is proven to be willful.
- CMI Penalties: Up to $25,000 per violation for the intentional removal or alteration of copyright management information.
- Injunctive Relief: A court-supervised order compelling Anthropic to destroy all infringing copies of their compositions.
- Transparency Metrics: A complete, audited accounting of Anthropic’s training data sources, training methodologies, and the specific lyrics utilized to construct its models.
Despite taking an aggressive stance against Anthropic, the publishers took care to position themselves as open to lawful technological innovation. The complaint notes that the plaintiffs recognize the creative potential of ethical AI and have actively established licensing agreements allowing authorized uses of their musical compositions with other AI developers.
“Even the most revolutionary of technologies must develop within the bounds of the law, and Anthropic’s Claude models are no different,” the complaint concludes.
Future Outlook: Industry Implications and Next Steps
As the legal battle unfolds in the U.S. District Court for the Northern District of California, the implications extend far beyond the offices of Sony Music Publishing, Warner Chappell Music, and Anthropic.
Legal analysts note that naming individual executives like Dario Amodei and Benjamin Mann marks a dangerous new precedent for tech founders. If the plaintiffs successfully establish personal liability for orchestrating data acquisition strategies, venture capitalists and startup founders across the tech sector may face unprecedented personal legal exposure for the data practices of the companies they build.
Furthermore, the reliance on unsealed evidence from the Bartz v. Anthropic settlement establishes a damaging narrative regarding corporate awareness. Proving willfulness in copyright litigation typically requires demonstrating that a defendant knew their actions were unlawful; internal Slack messages labeling source repositories as “sketchy AF” and expressing a desire for secrecy provide plaintiffs with potent ammunition.
As of the initial filing date, Anthropic has not yet issued a public response to the lawsuit. However, with heavyweights Oppenheim + Zebrak and Pryor Cashman representing the music publishers, the legal confrontation promises to set a definitive legal boundary defining how artificial intelligence companies source, ingest, and monetize creative content in the digital age.
