NEW YORK / RankWire.AI / – Three leading U.S. publishing companies have filed a lawsuit against Google, accusing the tech giant of copyright infringement related to its Gemini artificial intelligence platform. Hachette Book Group, Cengage Learning, and Elsevier initiated a proposed class action along with author Scott Turow and his firm, S.C.R.I.B.E. They submitted their complaint on July 10 in a federal court in New York. The suit claims Google unlawfully copied millions of copyrighted books and journal articles during the development and training of Gemini models.

The plaintiffs argue that Google sourced material via Google Books, Google Play Books, and Google Scholar. Publishers and authors had contributed works to support features like search, sales, and research services, the complaint states. However, the filing asserts that these arrangements did not authorize Google to reproduce the works for commercial AI training purposes. Additionally, Google is accused of utilizing web-scraped datasets that included content from pirate sites and subscription services behind paywalls.
Google faces four legal claims outlined in the 57-page complaint. Three of these claims allege illegal reproduction through Google services, web scraping activities, and the training or development of Gemini. The fourth claim involves violations of the Digital Millennium Copyright Act. The plaintiffs accuse Google of removing or altering copyright management information, including author names, ownership details, and publication data. As of July 15, the court had not yet ruled on the allegations or granted class-action certification.
Four allegations focus on Gemini’s training data
The proposed class includes owners of registered U.S. copyrights in books and journal articles. Eligible works must have an International Standard Book Number. Articles must possess a Digital Object Identifier or International Standard Serial Number. This definition applies to works Google allegedly copied from its services, retrieved via web scraping, or reproduced during the development of Gemini. Eligibility is also limited to works registered within the deadlines specified in the complaint.
The complaint cites works from Hachette, Cengage, and Elsevier as examples of alleged copying. It covers fiction, textbooks, and scholarly publications. The filing also references internal Google assessments concerning legal risks associated with publisher-supplied books. One such assessment reportedly warned of potential fines between $10 billion and $100 billion, according to the plaintiffs. The court has not issued any rulings regarding these internal documents.
Plaintiffs seek damages and transparency
The plaintiffs seek statutory or actual damages along with profits attributable to any proven infringement. They also request an injunction, coverage of legal expenses, and a jury trial. Their proposed order would require Google to disclose the materials used and methods employed in training Gemini. The complaint further calls for court supervision to ensure the destruction of any unauthorized copies under Google’s control. The total damages sought are not specified in the filing.
This New York case follows an earlier effort by Hachette and Cengage to incorporate claims into a separate Google AI copyright lawsuit in California. The Association of American Publishers noted that this new filing preserves certain claims outside the scope of that case. The current lawsuit involves Elsevier, Turow, and S.C.R.I.B.E., alongside the two publishers. It asks the New York court to determine whether Google’s Gemini training practices and data collection activities infringe upon federal copyright law and the Digital Millennium Copyright Act.