← Blog

Hachette v. Google: Inside the Publishers' Gemini AI Book-Copyright Lawsuit

September 24, 2026

On July 10, 2026, a group of publishers and authors — Hachette Book Group, Cengage Learning, Elsevier, bestselling novelist Scott Turow, and S.C.R.I.B.E., Inc. (the entity that holds the copyrights in Turow's novels) — filed a proposed class-action copyright suit against Google in the U.S. District Court for the Southern District of New York. The case is Hachette Book Group, Inc. v. Google LLC, No. 1:26-cv-05870. The plaintiffs allege that Google copied millions of copyrighted books and journal articles without permission to build and train its Gemini generative-AI models. The complaint opens with a line the coverage could not resist: "Desperate to maintain its online dominance, Google abandoned its early motto of 'Don't be evil' and engaged in one of the most prolific infringements of copyrighted materials in history." The filing was picked up across the trade and mainstream press, from Publishers Weekly and TheWrap to Al Jazeera and the Indian Express.

What makes this one worth a closer look isn't just the size of the numbers, though observers have floated exposure well into the tens of billions of dollars. It's the seam it exposes — one that runs directly under every creator who has ever handed their work to a platform. The plaintiffs aren't only complaining that Google scraped the open web. They allege that some of the very books at issue were given to Google under narrow, purpose-specific agreements — for Google Books, Google Play Books, and Google Scholar — and were then quietly repurposed to train a commercial AI that now competes with those same books. That is the part every filmmaker, author, and rights holder should sit with: the question isn't whether Google had the files. It's what Google was allowed to do with them.

What the lawsuit says

According to the complaint and the coverage around it, the plaintiffs bring four claims: direct copyright infringement, contributory copyright infringement, removal or alteration of copyright-management information (CMI), and related violations of the Digital Millennium Copyright Act. The core factual allegation is that Google assembled a vast corpus of copyrighted works from several channels — its long-running Google Books scanning program, publisher and author submissions made to Google Play Books and Google Scholar under limited-use agreements, web scraping, and what the plaintiffs describe as known pirate sources — and then reproduced those works many times over during the training of Gemini.

The plaintiffs frame the harm in market terms, which is where these AI cases increasingly live. They argue Google bypassed a real, emerging market in which AI developers pay publishers to license training content — choosing to take rather than license — and that Gemini now competes directly with the works it learned from by generating summaries, textbook-style explanations, and other written material that can substitute for the originals. On the DMCA theory, they allege Google stripped or altered the copyright-management information attached to the works, the digital identifiers that tell the world who owns what. The plaintiffs seek statutory damages, a permanent injunction, and destruction of the unauthorized copies. They also take pains to note this suit is distinct from an earlier 2023 California class action over Google's AI practices; this is a new, separately pleaded case.

Google has not yet answered on the merits, and nothing has been decided. These are allegations in a complaint, not findings; Google has not conceded that any of its conduct infringes anything, and being named as a defendant is not evidence of wrongdoing. Google has defended its AI training as lawful in other matters, and it will have every opportunity to do so here. What follows is the legal question the filing raises — not a prediction of how it resolves.

The hard part: transformative use, and the ghost of Google Books

There is a reason Google, of all defendants, draws this particular suit — and a reason the plaintiffs wrote their complaint the way they did. A decade ago, the Authors Guild sued Google over the same Books program, arguing that scanning millions of volumes to make them searchable was mass infringement. Google won. In 2015 the Second Circuit held that scanning books to build a search index and show short snippets was a transformative fair use, and the Supreme Court declined to take the case. That ruling is the backdrop to everything here, and both sides know it.

The plaintiffs' bet is that the 2015 win does not stretch to cover 2026. Indexing a book so a reader can find it, they will argue, is a fundamentally different act from ingesting that book to train a system that generates competing prose. Snippet search points people toward the original and can even sell copies; a generative model trained on the full text can produce output that stands in for the original and suppresses demand. Fair use turns heavily on the purpose of the use and its effect on the market for the work, and the plaintiffs have pleaded both to distinguish their case from the one Google won. Google, for its part, will lean on the transformative-use logic that carried Google Books, argue that training is a non-expressive statistical process, and dispute that Gemini's outputs meaningfully substitute for anyone's book. The limited-use-agreement allegations add a second front the older case didn't have: even if training were fair use in the abstract, taking works that were handed over for Books, Play, or Scholar and using them for something else is a contract and CMI problem, not just a fair-use debate. Where the record shows a clean, licensed-and-limited transfer that was exceeded, the plaintiffs' claims sharpen. Where it shows lawfully accessible material used transformatively, Google's defense widens.

Why it matters beyond one lawsuit

Strip away the household names and this is a lesson the IP Feed keeps circling back to: the permission you grant for one use is not a license for every use — especially not for AI training, a use most contracts written even a few years ago never contemplated. Publishers gave Google their books to be found. They did not, they say, agree to have them fed into a model that competes with them. That gap between "what I allowed" and "what was done" is exactly where creators get hurt, and it is not a book-industry problem. It is a filmmaker problem, a musician problem, a photographer problem. Every distribution deal, platform upload, and archive submission carries grant-of-rights language, and increasingly that language reaches for phrases like "new technologies now known or hereafter developed" — the clause that can quietly sweep AI training into a deal you thought was about streaming.

The practical takeaways are concrete. Read the grant-of-rights and license clauses in every distribution, platform, and licensing agreement, and look specifically for AI, machine-learning, and "future technologies" language — narrow it or carve it out before you sign, because a limited-use grant is only worth what the words say. Register your copyrights: statutory damages and the strongest remedies, including the DMCA and CMI claims at the heart of this case, depend on it, and registration is far cheaper than litigation. Keep your copyright-management information — credits, ownership metadata, watermarks — intact and attached to your files, because stripping it is itself a claim. And recognize that an AI-licensing market now exists; work that is being ingested for free today is work someone is willing to pay to license, which means your catalog has a value you can choose to sell rather than surrender. For creators who spend years making something worth copying, the difference between a taking and a license is usually a few sentences in a contract signed long before anyone thought about it.

We'll track this docket as it develops. The first real signals will come as Google answers — whether it moves to dismiss, argues that training Gemini is transformative fair use in the mold of Google Books, disputes the limited-use and CMI theories, or challenges class certification — and as the court begins sorting the copyright and DMCA claims. Follow the filings, counsel, and coverage on the case page.

This post is editorial commentary on public court filings and news coverage, not legal advice. The allegations described are unproven, the defendant has not yet responded on the merits, and being named as a defendant is not evidence of wrongdoing. Details are drawn from the docket and press reports and may be refined as the case develops.