← Blog

Sullivan v. OpenAI: Inside the Textbook Authors' AI-Training Copyright Class Action

September 30, 2026

On August 14, 2026, a group of textbook authors — the people whose names sit on the spines of the books that taught a generation of American students calculus, finance, and human anatomy — walked into the federal courthouse in Manhattan and sued the most closely watched company in artificial intelligence. The case is Sullivan v. OpenAI Foundation, No. 1:26-cv-06966, filed in the U.S. District Court for the Southern District of New York and assigned to Judge Sidney H. Stein. The lead plaintiff is Michael Sullivan, whose mathematics textbooks have been assigned in classrooms for decades. Standing with him are authors of some of the most widely used college texts in the country — including finance scholars Zvi Bodie and Alan J. Marcus, anatomy authors Kenneth S. Saladin and Kevin T. Patton (through his company Lion Den, Inc.), and business-communication author Dana Loewy. Their targets: OpenAI, the maker of ChatGPT, and its largest backer and commercial partner, Microsoft.

This is a putative class action — the seven named authors are asking the court to let them represent a far larger group, potentially thousands of textbook authors whose works they say met the same fate. And the fate they allege is blunt: their books were copied, without permission or payment, and poured into the training data behind OpenAI's large language models.

The allegation: shadow libraries as training fuel

At the center of the complaint is a claim that has become the recurring theme of the AI copyright wars — that the books were not licensed, purchased, or scraped from the open web, but pulled from pirated "shadow libraries." These are the sprawling, unauthorized repositories that circulate digitized copies of copyrighted books, and plaintiffs across a wave of AI cases have pointed to them as the original sin of model training. The authors contend that when their textbooks entered a training corpus assembled from those sources, OpenAI and Microsoft reproduced protected works at scale to build a commercial product — and did it knowing exactly whose work they were using.

As with every dispute we cover, these are allegations. OpenAI and Microsoft have not conceded the copying, dispute that any of it amounts to infringement, and have argued that training a model on text is a transformative, lawful use. No court has ruled on the merits.

The four claims, and the one that is easy to miss

The complaint stacks the copyright theories the way most of these cases do:

Then comes the claim that is easy to skate past but may matter most for creators. The authors also invoke Section 1202 of the Digital Millennium Copyright Act, which prohibits stripping or altering copyright management information — the author names, titles, ISBNs, and copyright notices attached to a work. The theory: when a book is ingested into a training set, the identifying metadata that says who made this and who owns it gets removed, so the model can learn from the text while the fingerprint of ownership disappears. For a textbook author, the copyright notice on page ii is not a formality. It is the machine-readable proof of authorship — and the DMCA claim is an argument that erasing it is itself a violation, separate from the copying.

Why textbooks are a sharp test case

Textbook authors make unusually strong copyright plaintiffs, and that is a large part of why this case is worth watching. Their works are registered. Their authorship is documented down to the edition and the printing. The chain of ownership between the author, the publisher, and the copyright office is about as clean as it gets in American publishing. When Michael Sullivan says a model trained on his mathematics text, he can point to a registration certificate, a publication date, and a title that has been in print through numbered editions for years.

That evidentiary cleanliness cuts to the heart of what these lawsuits are really fighting about. AI companies have argued that training is a transformative fair use, and that the sheer scale of the data washes out any single work's contribution. The authors are countering with specificity: named works, registered rights, and a documented before-and-after. The stronger the provenance, the harder it is for a defendant to wave the works away as an anonymous drop in an ocean of text.

One front in a much larger war

Sullivan does not stand alone. It has been folded into the consolidated proceeding In re: OpenAI, Inc. Copyright Infringement Litigation, the multidistrict docket before Judge Stein that has become the center of gravity for copyright claims against the company — pulling together authors, news organizations, and other rights holders whose cases raise overlapping questions about training data and fair use. That same litigation is where courts have begun the slow work of sorting which claims survive early motions and which get trimmed, testing each theory before any of it reaches a jury.

Regular readers will recognize the neighborhood. We have covered the Seattle Times and Newsday's suit against OpenAI and Microsoft over paywalled articles, and the parallel campaigns by book publishers and music companies against the other AI labs. The textbook authors are a distinct constituency inside that fight, but the core question is the same one running through all of it: when a machine learns from a copyrighted work, who owes whom, and for what?

The stakes rose further in September 2026, when the U.S. Department of Justice — listed on the docket as an interested party — weighed in with a statement expressing skepticism that training an AI model on copyrighted text is automatically infringement. That posture does not decide any case, but it signals how contested the ground remains, and how much every well-documented plaintiff matters to how the law ultimately lands.

Why this matters for rights holders and creators

Strip away the AI-training specifics and Sullivan v. OpenAI is a story about leverage — the leverage that comes from being able to prove, precisely, what you made and what you own. These authors can take on OpenAI and Microsoft not because they are large or powerful, but because their rights are registered, dated, and specific. They can name the work, produce the certificate, and point to the copyright notice that was supposed to travel with it. That documentation is what turns a grievance into a claim.

The lesson generalizes to every creator, and it is the same throughline that connects a calculus textbook to an independent filmmaker's rights ledger. The value of any asset — a book, a song, a screenplay, a finished feature — is only as strong as your ability to prove what it is, when it came into being, and exactly what rights you hold. A catalog that can show its provenance can go on offense: it can license, enforce, or demand payment when someone uses it without permission. One that cannot is left hoping no one ever tests it. The textbook authors kept their registrations clean, and that record is now the foundation of a case against two of the most formidable defendants in technology.

For filmmakers and rights holders, the takeaway is not that they will someday sue an AI lab. It is that a provable chain of title — clear authorship, registered rights, documented ownership — is not paperwork for its own sake. It is the thing that lets you defend, license, or monetize what you created, whoever ends up using it. In a world where copying has never been easier, the party with the cleaner record of ownership sets the terms.

We will track the consolidated docket as it develops — how the court treats the direct-infringement and DMCA claims, whether the textbook-author class is certified, and how the broader fair-use question resolves across the OpenAI cases. Follow the filings, counsel, and coverage on the case page.

This post is editorial commentary on public court filings, not legal advice. The infringement and DMCA claims described here are allegations that no court has adjudicated; the defendants dispute them and have raised fair-use and other defenses. Party names, docket details, dates, and claims reflect the public record as we found it. Nothing here should be read as a prediction of how the case will be resolved.