The newest publisher lawsuit against OpenAI and Microsoft is about copyright, but its most consequential argument may be about something broader: what happens when an AI answer becomes good enough that the reader never visits the publisher that produced the underlying journalism.
The Seattle Times and Newsday have sued OpenAI and Microsoft in federal court, alleging that the technology companies copied their journalism without authorization and used it to train and operate artificial intelligence products including ChatGPT, Microsoft Copilot and Bing’s AI features. According to Reuters’ September 5 report, the complaint says the alleged scraping included articles protected by paywalls.
The newspapers go further than arguing that their work entered an AI training pipeline without permission. They allege that AI products can reproduce passages, closely paraphrase their reporting and answer users’ questions in ways that reduce the need to visit the original websites or purchase subscriptions.
That claim gets to the heart of the conflict between publishers and generative AI. The dispute is no longer only about whether copyrighted material can lawfully be used to build a model. Publishers increasingly argue that the resulting products can compete with the very websites whose content helped make those answers possible.
The lawsuit targets both training and AI products in use
The case was filed in the U.S. District Court for the Southern District of New York. The public court docket for The Seattle Times Company et al v. OpenAI Inc. et al lists The Seattle Times Company and Newsday LLC as plaintiffs and multiple OpenAI entities plus Microsoft Corporation as defendants.
Reuters reports that the newspapers accuse the companies of scraping their websites and incorporating articles into datasets used both to train and operate AI systems. That distinction matters because copyright disputes around generative AI increasingly involve several stages: acquiring copyrighted material, using it during training, retrieving or processing it later and producing outputs that may resemble or summarize the original.
The publishers are seeking an order requiring the destruction of unauthorized copies of their works as well as training datasets or AI models that incorporate them, according to Reuters.
None of those allegations has been established by a court. The case is at the complaint stage, and OpenAI and Microsoft dispute the broader premise that their AI practices amount to unlawful copyright infringement.
The paywall allegation raises the stakes
The complaint’s allegation that paywalled journalism was scraped is particularly sensitive for subscription publishers.
A paywall is not merely a technical barrier. It is part of the publisher’s business model. The publication invests in reporting and asks readers to pay for access to some or all of that work.
If a third-party AI system can provide a detailed substitute for that reporting without requiring the user to subscribe, the publisher can argue that the harm is not limited to unauthorized copying. The AI product may interfere with the transaction the paywall was designed to create.
That is one reason the dispute over AI and journalism is different from a purely academic debate about training data. News organizations depend on a recurring economic cycle: reporting produces valuable information, readers visit or subscribe, and that revenue finances more reporting.
The publishers’ theory is that generative AI can interrupt that cycle by extracting information upstream and satisfying the user downstream.
The core allegation is substitution
Reuters summarizes one of the newspapers’ central claims clearly: the AI products can reproduce passages or closely paraphrase articles while giving users answers that reduce the need to visit the publishers’ websites or purchase subscriptions.
That is a substitution argument.
Traditional search engines also extract information from publisher pages, but their conventional model is built around sending users onward through links. Publishers have spent years accepting the tradeoff because search visibility could generate substantial referral traffic.
Generative interfaces change that bargain. A user can ask a question and receive a synthesized answer directly in ChatGPT, Copilot or an AI-enhanced search interface. A citation may be present, but the user may already have received enough information to stop.
For publishers, the economic question becomes whether being a source for the answer generates sufficient value to compensate for the visits that the answer potentially eliminates.
A citation does not necessarily equal a click
The AI industry has increasingly emphasized citations and links as a way to connect generated answers back to original sources. Those features matter, particularly for attribution and discoverability, but they do not fully answer publishers’ concerns.
A citation creates an opportunity to visit a source. It does not require the user to do so.
If an AI answer extracts the central facts, context and conclusion from an article, a user may have little reason to open the citation. This is especially significant for publishers whose business depends on advertising impressions, registrations or subscription conversion after the click.
The difference between being cited and being visited is becoming one of the most important measurement problems in AI search. A publisher can become highly visible inside generated answers while receiving comparatively little referral traffic from that visibility.
The Seattle Times and Newsday lawsuit pushes that issue from analytics into law by alleging that AI-generated answers can act as substitutes for the publishers’ own products.
OpenAI says its training practices are grounded in fair use
OpenAI did not specifically address the new lawsuit’s individual allegations in the statement reported by Reuters. A company spokesperson said OpenAI’s models are trained on publicly available data and that its practices are grounded in fair use.
Fair use is central to the U.S. legal battle over AI training. Technology companies have argued that training models on copyrighted material can be transformative and legally permissible, while publishers and other copyright holders argue that copying their works without permission—particularly to build commercial products that may compete with them—violates their rights.
Courts are still defining where those boundaries lie. The legal analysis can depend on the specific works, how they were acquired, what the model does with them and whether the resulting use affects the market for the originals.
The Seattle Times and Newsday case joins a much larger wave of litigation rather than resolving that debate.
Microsoft says it is surprised and open to discussion
Microsoft told Reuters that it was surprised by the lawsuit, while emphasizing the importance of local journalism and saying it was willing to sit down and explore solutions to this kind of dispute.
That response highlights an awkward aspect of the relationship between technology companies and news organizations. AI companies need large amounts of high-quality human-created information, and professional journalism is particularly valuable because it contains original reporting, editing, verification and current facts.
At the same time, publishers increasingly see AI answer engines as potential competitors for reader attention.
The parties therefore have incentives both to cooperate and to fight. Licensing deals can create revenue and clearer permissions, while litigation can establish boundaries for content that was allegedly used without such agreements.
The case echoes The New York Times lawsuit
The new complaint follows the landmark lawsuit The New York Times filed against OpenAI and Microsoft in 2023. Reuters notes that the Times case remains ongoing and is one of dozens brought by copyright holders against AI companies including OpenAI, Anthropic and Meta.
The similarities are significant. Publisher lawsuits increasingly combine allegations about unauthorized copying with claims that generative products can reproduce, summarize or compete with copyrighted works.
What is evolving is the surrounding market. In 2023, ChatGPT was still a relatively new consumer behavior. By 2026, conversational AI has become embedded across search engines, browsers, operating systems and productivity software.
That expansion makes the substitution question more commercially important. The more users rely on AI interfaces as their first stop for information, the more publishers need to know whether those interfaces create incremental discovery or replace direct consumption.
Local journalism makes the economic argument especially sharp
The Seattle Times and Newsday are not abstract content databases. They are news organizations that spend money to report on communities, institutions and events.
Seattle Times President and CEO Alan Fisco told employees that the organization spends millions of dollars annually producing its content and believes it must defend that work from use without consent or compensation, according to Reuters.
That economics matters because original reporting has substantial fixed costs. A reporter must investigate the story whether one thousand or one million people eventually read it. Digital distribution is cheap; producing the information is not.
If AI systems can absorb the output of that investment and deliver a useful substitute at near-zero marginal cost, publishers argue that the market can become structurally imbalanced.
The technology company benefits from the information while the organization that paid to discover it risks losing the audience relationship that finances the next story.
The fight is moving from “Can AI train on this?” to “What does AI replace?”
Much of the first generation of AI copyright litigation focused on the training process. Did a company copy protected works? Was that copying authorized? Does fair use permit it?
Those questions remain central, but the publisher cases increasingly add another layer: market substitution.
If a model learns from an article but never exposes any recognizable expression and never affects demand for the original, the legal and economic arguments look one way. If a product can reproduce passages, closely paraphrase the article or provide an answer that satisfies the same demand as the original publication, publishers argue that the analysis should look very different.
This distinction is likely to become more important as AI products evolve from general-purpose chatbots into research assistants, search engines and personalized news interfaces.
Publishers need to measure AI visibility and AI displacement separately
For digital publishers, the lawsuit also exposes a measurement problem that exists regardless of how the court rules.
AI visibility can be positive. Being cited by ChatGPT or Copilot may introduce a publication to users who would never have discovered it through traditional search. A strong citation can reinforce authority and sometimes generate referral traffic.
But visibility and displacement are different metrics.
Publishers need to measure how often AI systems mention or cite them, how much referral traffic those citations produce, whether visitors convert into subscribers and whether search or direct traffic changes as answer engines become more capable.
A high citation count should not automatically be treated as a success if the same interface is satisfying the user’s information need before the visit.
Robots controls and licensing cannot solve every historical dispute
Publishers now have more tools for expressing preferences about AI crawlers than they did during the earliest wave of model training. Some AI companies also offer licensing agreements with news organizations.
Those developments may reduce future ambiguity, but they do not automatically resolve claims about content allegedly acquired or used in the past.
Nor do they answer the broader competitive question. A publisher may allow an AI system to access content because it wants visibility, while still worrying that the resulting answers reduce traffic. Conversely, blocking access may protect content while removing the publication from an increasingly important discovery channel.
The business decision is becoming as complicated as the copyright decision.
The future of AI search depends on whether the source ecosystem can survive the answer
The Seattle Times and Newsday lawsuit presents its claims in forceful terms, but the underlying economic question extends far beyond these two newspapers.
Generative AI systems become more useful when they can draw on accurate, current, professionally produced information. Journalism becomes harder to finance when readers no longer need to visit or subscribe to the organizations producing that information.
Those two trends can coexist for a while. They are harder to sustain indefinitely if the value flows predominantly in one direction.
OpenAI and Microsoft will have the opportunity to contest the newspapers’ allegations, and courts—not publishers—will decide the legal merits. But the lawsuit identifies a question that every AI answer engine will eventually have to confront.
The issue is not only whether an AI system used a publisher’s content. It is whether the answer it created became a replacement for the visit that was supposed to pay for the journalism in the first place.