Encyclopedia Britannica, along with its affiliate Merriam-Webster, has initiated legal proceedings against OpenAI, asserting in their complaint that the artificial intelligence powerhouse is guilty of "massive copyright infringement."
Britannica, the parent company of Merriam-Webster, claims ownership of the copyrights for close to 100,000 online articles. The publisher contends in the lawsuit that these copyrighted works have been extensively scraped and utilized without authorization to train OpenAI’s large language models (LLMs).
Furthermore, Britannica accuses OpenAI of infringing copyright legislation through the generation of outputs featuring “full or partial verbatim reproductions” of its proprietary content. The lawsuit also highlights the AI lab's use of its articles within ChatGPT’s Retrieval Augmented Generation (RAG) workflow – a mechanism through which the LLM accesses current information from the web or other databases to formulate responses. Adding another layer to its claims, Britannica alleges a violation of the Lanham Act, a trademark statute, stemming from OpenAI's generation of fabricated "hallucinations" that are then falsely attributed to the publisher.
The complaint states that “ChatGPT starves web publishers like [Britannica] of revenue by generating responses to users’ queries that substitute, and directly compete with, the content from publishers like [Britannica].” Beyond financial harm, Britannica asserts that ChatGPT’s propensity for generating "hallucinations" endangers “the public’s continued access to high-quality and trustworthy online information.”
Britannica's legal action places it among a growing cohort of publishers and authors who have initiated lawsuits against OpenAI concerning copyright infringements. Prominent plaintiffs include The New York Times, Ziff Davis (which owns Mashable, CNET, IGN, and PC Mag, among others), and over a dozen newspapers spanning the U.S. and Canada, such as the Chicago Tribune, the Denver Post, the Sun-Sentinel, the Toronto Star, and the Canadian Broadcasting Corporation.
A comparable lawsuit filed by Britannica against Perplexity remains ongoing.
The legal landscape regarding whether the use of copyrighted material for training large language models constitutes infringement lacks clear precedent. However, a notable case involving Anthropic saw federal judge William Alsup rule that this specific use — employing content as training data — could be deemed transformative and thus legal. Nonetheless, Judge Alsup found Anthropic in violation of the law for illicitly downloading millions of books rather than licensing them, leading to a substantial $1.5 billion class action settlement for affected writers.
OpenAI did not provide a response to TechCrunch’s inquiry for comment prior to the article’s publication.
The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.