Hachette, Elsevier, Cengage Learning and author Scott Turow have sued Google in Manhattan federal court, alleging Gemini was trained on millions of books obtained through snippet-only agreements — the fourth major AI copyright suit filed by the same publisher coalition and the first to target a company that, unlike Meta, Microsoft, Amazon, OpenAI and Anthropic, has struck no publisher licensing deals.
For three years, publishers have tried nearly everything short of a lawsuit to get AI companies to pay for the books powering their chatbots: crawler blocks, public pressure, licensing negotiations, and — most recently — Cloudflare’s plan to block multi-purpose scrapers by default starting September 15, a move aimed squarely at Google’s dual-use crawler, according to ADWEEK. None of it brought Google to the negotiating table. Now the publishers are trying court.
Hachette Book Group, Cengage Learning, Elsevier and bestselling author Scott Turow filed a proposed class action against Google on July 10 in the U.S. District Court for the Southern District of New York, accusing the company of illegally copying millions of books and journal articles to train its Gemini AI models, as Publishing Perspectives reported. The plaintiffs — among the largest names in trade, academic and scientific publishing — are seeking statutory damages, an injunction, and an order forcing Google to destroy unauthorized copies, according to TheWrap.
A promise repurposed
The suit’s central claim is not that Google scraped random web pages, but that it took books submitted for narrow, specific purposes and fed them into a general-purpose AI. Publishers gave Google access to books for Google Books, Google Play and Google Scholar under agreements that, the complaint says, permitted only search indexing, snippet display, or ebook retail — not training a model built to compete with the books themselves, according to Startup Fortune. Publishing Perspectives noted the complaint leans on a 2015 appeals court ruling that found Google’s book-scanning project was fair use because it was “highly transformative” — while acknowledging that ruling never explicitly extended to AI training, a gap the publishers argue Google exploited.
the paper trail
The complaint’s sharpest allegation rests on Google’s own internal assessment of its risk. According to reporting cited by Startup Fortune, internal Google documents flagged book training as “highly problematic” and estimated potential copyright exposure at “$10Bs-$100Bs” — a figure ADWEEK also reported from the complaint. The suit further alleges Google stripped or altered copyright management information to obscure the origin of its training data. The plaintiffs argue the numbers show a company that ran the math and decided litigation risk was worth taking rather than paying for licenses.
The complaint also points to concrete competitive harm: Gemini, the filing states, can generate “a 100-page murder mystery set in a quiet seaside town filled with secrets” that substitutes for an original copyrighted novel in “20 minutes” for 39 cents, according to Publishing Perspectives. “No publisher or author can compete with that,” the filing states.
a familiar plaintiff roster
This is not the coalition’s first swing. In May, a nearly identical group — Elsevier, Cengage and Hachette, joined by Macmillan and McGraw Hill, with Turow again attached — sued Meta in the same Manhattan courthouse over Llama’s training data, ADWEEK reported. Meta has denied wrongdoing and argues training on copyrighted material can qualify as fair use. The strategy has paid off at least once: Anthropic agreed last year to pay authors $1.5 billion to settle a similar piracy claim, a deal Judge William Alsup gave preliminary approval in September 2025 and under which the company pays roughly $3,000 for each of about 465,000 books, according to the Associated Press as cited by Startup Fortune. The New York Times’ parallel suit against OpenAI and Microsoft remains ongoing. Alsup’s earlier ruling drew a distinction the Google case will test further: training on lawfully acquired books can be fair use, but building a library from pirated copies cannot — leaving open how courts will treat books obtained under narrow, purpose-limited contracts that were neither pirated nor licensed for AI.
google’s licensing gap
Unlike Meta, Microsoft, Amazon, OpenAI and Anthropic, Google has struck no licensing agreements with digital publishers, ADWEEK has previously reported — though it does have one with the Associated Press, a wire service. That refusal has pushed media companies toward what ADWEEK called “once-unthinkable measures,” including USA Today Inc. CEO Mike Reed’s statement that his company is prepared to delist from Google Search within six to 12 months absent a licensing deal. Google did not immediately respond to a request for comment on the lawsuit, according to TheWrap, and reporters at TechCrunch and Publishing Perspectives both noted the company had no comment on the specific allegations as of this week.
What happens next will hinge on how a court treats content handed over for one narrow purpose and later repurposed for a far more valuable one. If the publishers’ “scope-limited” theory holds, it could reshape how every tech company treats archives it built for search, retail or indexing — not just Google’s. Given the Anthropic settlement and the unresolved Meta and OpenAI cases, Google now faces the added burden of its own internal memo putting a price on the risk it appears to have taken anyway.
Hatchette and Elsevier Sue Google for Using Their Work to Train AI
— Gizmodo