NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning and Elsevier have initiated a lawsuit against Google regarding its Gemini artificial intelligence platform. Author Scott Turow along with his company, S.C.R.I.B.E., have joined the proposed class action. The legal complaint was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without permission during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or approve the class action certification.

According to the complaint, Google acquired material through Google Books, Google Play Books, and Google Scholar. Publishers and authors provided works for specific uses such as search, sales, and research purposes. The plaintiffs argue that these agreements did not authorize the broader use of their works for commercial AI training. They also claim Google downloaded extensive web-scraped datasets containing copyrighted content, some of which originated from known pirate sources and paywall-protected services.
The 57-page complaint outlines four claims under federal law. Three of these concern alleged reproduction through Google services, web scraping activities, and the development or training of Gemini. The fourth is based on the Digital Millennium Copyright Act. The plaintiffs allege Google removed or altered copyright management information from training data. The document also references internal discussions about using publisher-provided books, with one assessment estimating potential fines ranging from $10 billion to $100 billion. The court has not yet examined these allegations.
Class Definition Includes Registered Works
The proposed class encompasses owners of registered U.S. copyrights in qualifying books and journal articles. To qualify, books must have an International Standard Book Number (ISBN), while articles need a Digital Object Identifier or International Standard Serial Number. The class includes works allegedly copied from Google services or obtained through web scraping, as well as those reproduced during Gemini’s development or training phases.
Eligibility for class membership also depends on registration timing. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another stipulates registration within three months of publication. The complaint excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before the case advances on behalf of the larger group.
Claims for Damages and Financial Accountability
The plaintiffs seek either statutory damages or actual damages related to proven infringements. They also request Google’s profits attributable to any confirmed copyright violations. Their requested remedies include an injunction, legal costs, and a jury trial. The complaint does not specify a total damages amount but urges Google to disclose Gemini training data, collection methods, and known capabilities through a court-mandated accounting.
This accounting would identify the copyrighted works used in training Gemini and detail how Google collected, copied, processed, and encoded these materials. The plaintiffs also request that the court oversee the destruction of any unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate legal actions against Google’s AI initiatives in California. This New York case expands the list to include Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google services, web scraping, and Gemini’s training process.
