← Blog

Hachette v. Google: Inside the Publishers' Gemini AI Book-Copyright Lawsuit

September 4, 2026

On July 10, 2026, a group of publishers and authors led by Hachette Book Group filed a federal copyright lawsuit against Google, alleging that the company copied millions of copyrighted books without permission to train its Gemini AI models. The case is Hachette Book Group, Inc. v. Google LLC, No. 1:26-cv-05869, filed in the Southern District of New York and assigned to Judge Loretta A. Preska, under the Copyright Act (17 U.S.C. § 101 et seq.). With nearly 80 press articles across some 70 outlets tracked in its first weeks, it is one of the most heavily covered new copyright filings on the IP Feed.

It is also a case with a long memory. One of the named plaintiffs is Scott Turow — the novelist and former Authors Guild president who helped lead the original fight over Google Books two decades ago. The plaintiffs' roster reads like a cross-section of the publishing industry: Hachette, the academic-and-reference houses Elsevier and Cengage Learning, and Turow himself, filing on behalf of a proposed class. The question they are putting back to the same court is a familiar one asked in an unfamiliar context: what is Google allowed to do with a library of books it never licensed?

What the lawsuit says

The claim is copyright infringement. According to the complaint as described across widely reported accounts, Google copied and ingested millions of copyrighted books — and, per some reporting, journal articles and other written works — to build the training corpus behind Gemini, its flagship generative-AI system. The plaintiffs say that copying was done without a license, without permission, and without payment, and that it lies at the root of a product Google now sells and monetizes.

Framed as a proposed class action, the suit seeks to speak for a much larger population of authors and publishers whose works, the plaintiffs allege, were swept into the same training pipeline. The remedies at stake in a case like this are the ones that make copyright owners' lawyers pay attention: statutory damages, which do not require proving specific financial harm and which scale with the number of works infringed, plus injunctive relief that could reach how the models were built and how they are used going forward.

Google has not yet answered the complaint in court, and the allegations are unproven. Historically, the company's position in book-copying disputes has rested on fair use — the argument that what it did with the text was transformative enough to fall outside the copyright owner's control.

Why this one rhymes with history

For anyone who followed the first Google Books war, the déjà vu is the point. Beginning in 2005, the Authors Guild and a group of publishers sued Google over its project to scan and index the world's books. That case ran for a decade and ended in 2015, when the Second Circuit held that Google's scanning — to power a searchable index that returned only short snippets — was fair use. The Supreme Court declined to take it up. Google won.

This filing lands in the same courthouse, with one of the same faces on the plaintiff side, but the technology has moved the ground underneath the old ruling. The 2015 decision was about search: copying books so people could find them, while showing users only fragments. The new complaint is about generation: copying books so a model can produce new text on demand. Whether the reasoning that protected a snippet-returning search index also protects a system designed to absorb and re-express what it read is precisely the open question — and it is not one the earlier case answered.

That distinction — indexing versus generating — is the fault line running through nearly every AI-training copyright fight now on the docket, from the author and music-publisher suits against Anthropic to the visual-art claims against image generators. Coverage of the Google case frequently invokes those parallel fights, including the recently approved Anthropic authors' settlement reported at roughly $3,000 per work, as a rough marker of what exposure at book-corpus scale could look like. Those comparisons are context, not precedent: no court has yet ruled on the merits of whether training a generative model on copyrighted books is fair use.

Why it matters beyond one lawsuit

Books are, in one sense, the cleanest possible battleground for this question. Unlike loosely tracked web text, published books come with registrations, ISBNs, contracts, and publishers who know exactly what they own and who have spent a century licensing it. "You should have licensed this" is not a hypothetical argument to a house like Elsevier or Cengage — licensing is the business. That makes provenance, not philosophy, the likely center of gravity: what was copied, where the files came from, whether any of it was licensed, and what happened to the identifying information attached to each work.

For anyone who creates or owns rights — a novelist, a documentary producer, a studio sitting on a library — the lesson is the same one these cases keep teaching. The value of a creative work is inseparable from a clean, provable record of who owns it and how it has been used. When a dispute finally arrives, that paper trail is the asset. Rights that are well-documented can be defended and licensed; rights that are fuzzy tend to become someone else's training data.

We'll track the docket as it develops — the first real signal will come when Google responds and the shape of its defense (transformative use, disputed sourcing, or both) comes into view, and whether the court treats the 2015 Google Books ruling as a shield or as a case about a different technology entirely. Follow the filings, counsel, and coverage on the case page.

This post is editorial commentary on public court filings and news coverage, not legal advice. The allegations described are unproven, and the defendants have not yet responded in court.