Publishers Say Bing AI and ChatGPT Don’t Just Use Their News — They Replace the Visit

Publishers Say Bing AI and ChatGPT Don’t Just Use Their News — They Replace the Visit
Sponsored

The Seattle Times and Newsday have taken one of the central arguments in the AI-publishing conflict to federal court: the problem, they say, is not merely that artificial intelligence systems learned from journalism without permission. It is that the resulting products can answer the reader's question so completely that the reader no longer needs to visit the newspaper that paid to report the story.

In a lawsuit reported by Reuters on September 5, the two newspapers accuse OpenAI and Microsoft of copying their journalism without authorization to train and operate AI systems. The complaint alleges that scraping extended to material behind paywalls and that the resulting content was incorporated into datasets connected with products including ChatGPT, Microsoft Copilot and Bing's AI features.

The publishers further allege that those systems can reproduce passages from their reporting, closely paraphrase articles and provide answers that reduce the need for users to visit their websites or purchase subscriptions. Those are allegations in a newly filed lawsuit, not established findings of infringement or proven measurements of lost traffic.

But the theory of harm is important because it links two debates that have often been treated separately: copyright at the input stage and traffic substitution at the output stage.

The lawsuit alleges use of paywalled journalism

According to Reuters, the complaint filed in the U.S. District Court for the Southern District of New York alleges that OpenAI and Microsoft scraped the newspapers' websites, including content behind paywalls.

That allegation raises a more pointed issue than the broad question of whether publicly accessible web material can be used to train AI models under fair use. Subscription journalism is intentionally placed behind a commercial access boundary. Publishers charge readers precisely because the reporting has economic value that they do not make freely available to every visitor.

The plaintiffs are effectively arguing that an AI company should not be able to bypass that business model, ingest the protected work and then produce a competing informational experience derived from it.

Whether the alleged acquisition occurred as described, which works were used and whether any such uses were legally permissible will be questions for the litigation.

The case targets both training and operation

The complaint, as described by Reuters, is not limited to historical model training. It alleges that newspaper articles were incorporated into datasets used to train and operate AI products.

That distinction matters because generative systems increasingly combine pretrained models with retrieval, search indexes and other sources of current information. Copyright disputes therefore no longer fit neatly into a single question about what went into a model years ago.

Publishers can challenge the acquisition of training material, the retrieval of current pages, the reproduction of protected expression in an answer and the commercial effect of the resulting interface.

The Seattle Times and Newsday case places several of those concerns inside the same complaint.

The publishers say AI answers can substitute for the article

The most consequential allegation for search and publishing strategy is that AI products do not merely point users toward journalism. They can substitute for the visit.

Reuters reports that the newspapers say ChatGPT, Copilot and Bing AI features can reproduce passages, closely paraphrase articles and answer users in ways that reduce the need to open the original site or buy a subscription.

This is the economic bridge between copyright law and zero-click search. A publisher can theoretically receive attribution and still lose the user if the answer contains enough information to satisfy the underlying need.

From the publisher's perspective, the question is therefore not only “Was our reporting used?” but “Was our reporting used to build an experience that competes with the page and subscription that financed it?”

Attribution does not necessarily restore the visit

AI companies have increasingly emphasized links and citations as a way to support publishers and make generated answers verifiable. Those features are meaningful because they can expose sources to users who might otherwise never encounter them.

But a citation and a click are different events. A user can read an AI-generated summary, see that it is supported by a newspaper and still decide that opening the underlying article is unnecessary.

This creates a difficult value-exchange problem. The source can gain visibility and authority while losing the measurable referral that historically converted search discovery into advertising impressions, registrations or subscriptions.

For subscription publishers, the tension is even sharper when the answer allegedly conveys information that would otherwise require paid access.

OpenAI says its approach is grounded in fair use

OpenAI disputed the broader premise behind the lawsuit without commenting specifically on its individual allegations. A spokesperson told Reuters that the company's models are trained on publicly available data and that its practices are grounded in fair use.

Fair use is one of the central legal defenses in U.S. generative-AI copyright litigation. The analysis is fact-specific and considers factors including the purpose and character of the use, the nature of the copyrighted work, the amount used and the effect on the potential market.

AI companies have argued that model training is transformative and does not simply store and redistribute works as conventional copies. Copyright owners have countered that the systems were built through unauthorized copying and can generate material that competes with the original market.

The Seattle Times and Newsday allegations have not yet been adjudicated, so neither side's legal characterization should be treated as settled law.

Microsoft says it was surprised by the lawsuit

Microsoft gave Reuters a different response. The company said it was surprised by the filing, stressed that it appreciates the importance of local journalism and said it is willing to sit down and explore solutions to the dispute.

Microsoft's inclusion is significant because the allegations extend beyond ChatGPT itself. Copilot and Bing's AI experiences place generative answers inside Microsoft's own consumer discovery ecosystem.

The case therefore targets the relationship between the model provider and a major technology partner that distributes AI-generated information through search and assistant products.

That structure is likely to remain important as copyright litigation tries to assign responsibility across model developers, search platforms, cloud providers and product operators.

The newspapers are seeking destruction of datasets and models

The requested relief is unusually consequential. Reuters reports that The Seattle Times and Newsday are seeking an order requiring destruction of copies of their works as well as training datasets or AI models incorporating them.

Courts have not established that such sweeping relief is appropriate for the allegations in this case. But the request illustrates why training-data litigation carries risks beyond ordinary monetary damages.

If a court were ultimately to require deletion or destruction tied to copyrighted works embedded in a large training pipeline, implementation could raise difficult technical questions about provenance, dataset lineage and whether specific training inputs can be meaningfully removed from a deployed model.

For AI companies, data governance is therefore becoming part of litigation readiness rather than merely a research operations issue.

The lawsuit echoes The New York Times case

Reuters notes that the new action follows the major copyright lawsuit filed by The New York Times against OpenAI and Microsoft in 2023. That litigation remains ongoing.

The broader field has expanded dramatically since then. Authors, news organizations and other rights holders have brought dozens of cases against companies including OpenAI, Anthropic and Meta, testing how existing copyright doctrine applies to generative model development.

The Seattle Times and Newsday complaint does not automatically rise or fall with those other cases. Different plaintiffs can allege different acquisition methods, works, outputs and market effects.

But every new publisher case adds pressure for courts to clarify where large-scale machine learning fits within copyright's existing framework.

Local journalism makes the substitution argument especially sensitive

The Seattle Times describes itself as one of the relatively few independent, locally owned news organizations remaining in the United States. Newsday's core audience is similarly rooted in Long Island and the New York metropolitan region.

Local reporting has an economic structure that differs from generic web content. Reporters attend meetings, cultivate sources, file records requests and investigate institutions whose activities may have little national search demand.

The resulting article can then become the authoritative source from which summaries, aggregators and AI systems learn what happened.

If the original reporting bears the cost while downstream interfaces capture the user's attention, the sustainability question becomes more than a conventional dispute over referral percentages.

Both publishers have also participated in AI journalism initiatives

There is an important complication in the relationship between the parties. The Seattle Times and Newsday were both participants in the Lenfest Institute AI Collaborative and Fellowship program supported by OpenAI and Microsoft.

In OpenAI's announcement of the initiative, the company said Newsday would explore AI tools for public-data summarization and aggregation, while The Seattle Times would use AI platforms in advertising go-to-market work, sales training and analytics.

OpenAI and Microsoft each committed $2.5 million in direct funding and $2.5 million in software and enterprise credits to the broader two-year pilot, for up to $10 million combined across the initiative.

That cooperation does not resolve the present copyright allegations. It demonstrates something more interesting: publishers can simultaneously see AI as a useful newsroom and business technology while objecting to how their journalism is allegedly acquired and repurposed by external AI products.

The dispute is not “publishers versus AI”

That distinction is important because the industry debate is often flattened into two camps: publishers resisting technology and AI companies building the future.

In reality, many publishers are adopting AI internally for transcription, translation, research, archives, personalization and commercial workflows while demanding licensing, attribution or stronger control over how external models use their reporting.

The disagreement is therefore often about terms, permission and value exchange rather than whether artificial intelligence should exist in journalism at all.

A newspaper can believe AI improves its newsroom and still believe an external company infringed its copyright.

The legal issue and the traffic issue are related but not identical

Publishers should also avoid treating a decline in referral traffic as automatic proof of copyright infringement. Copyright law and web analytics answer different questions.

An AI answer could reduce clicks without infringing copyright if it independently provides facts or otherwise operates within lawful boundaries. Conversely, an infringing use could theoretically occur even if the publisher cannot demonstrate a measurable traffic decline.

The Seattle Times and Newsday complaint links the two by alleging both unauthorized copying and market substitution. The plaintiffs will still need to prove the legal elements of their claims and any damages they seek.

For SEO and publishing teams, however, the traffic mechanism matters regardless of how the copyright claim is ultimately resolved.

Search is becoming a destination rather than a routing layer

The lawsuit arrives as search engines increasingly answer questions inside their own interfaces. Traditional blue-link search was never a pure referral service, but the user's path commonly continued from the result page to a publisher.

Generative search can complete far more of the information task before that transition occurs. A sufficiently comprehensive answer can combine several sources, explain the event and answer follow-up questions without requiring the user to leave the platform.

This shifts the publisher's role from destination toward upstream information supplier.

That can be valuable when citations introduce new audiences. It can be economically destructive when the upstream source becomes essential to the answer but optional to the user.

Paywalls make the replacement question harder

A paywall is a market signal as much as a technical barrier. It says that the publisher intends to exchange access to this reporting for money, registration or another form of subscriber relationship.

If an AI system can provide the substance of that reporting outside the paywall, the publisher can argue that the system is not simply indexing the open web. It is interfering with the mechanism through which the work is sold.

The factual details matter enormously. A short factual summary of a news event is not equivalent to reproducing protected expression from an article, and copyright generally does not protect facts themselves.

The litigation will therefore need to distinguish between conveying facts, paraphrasing protected expression, reproducing passages and using works in model training.

The fight is ultimately about who captures the value after reporting happens

News organizations have always operated in an ecosystem of intermediaries. Search engines, social networks, aggregators and news apps have all shaped how audiences discover journalism.

Generative AI changes the bargaining dynamic because the intermediary can increasingly produce the final informational experience itself.

If a user asks what happened at a local council meeting and receives a complete answer synthesized from a newspaper reporter's work, the publisher may have funded the expensive part of the process while the assistant owns the user interaction.

The Seattle Times and Newsday are asking a court to decide whether the way OpenAI and Microsoft allegedly built and operate those systems violates copyright. Publishers across the industry are simultaneously asking a business question that courts cannot solve alone: what is a sustainable exchange when AI products can use journalism to satisfy the reader without delivering the reader?

For publishers, AI visibility may carry a hidden tradeoff

SEO teams increasingly measure whether their publications are cited in ChatGPT, Copilot, AI Overviews and other answer engines. That visibility can strengthen authority and expose reporting to audiences beyond conventional search.

But citation success should not automatically be interpreted as distribution success. Publishers need to measure whether AI visibility produces referral visits, branded demand, registrations, subscriptions or other value that compensates for potential click substitution.

A source can become more important to the information ecosystem while becoming less visited by the people consuming that information.

That is the paradox at the center of the new lawsuit.

The allegations now have to survive the legal process

The Seattle Times and Newsday have made serious claims: unauthorized copying, use of paywalled journalism, reproduction or close paraphrasing of reporting and market harm through answers that allegedly reduce visits and subscriptions.

OpenAI says its training practices rely on publicly available information and fair use. Microsoft says it values local journalism and is open to discussing solutions. The court has not established which account is legally correct.

Until evidence is tested, the case should be described as an allegation rather than proof that ChatGPT, Copilot or Bing unlawfully copied the newspapers or caused specific subscriber losses.

What is already clear is why the dispute matters beyond these two publishers. The web's old exchange was imperfect but understandable: publishers created information, search engines helped people find it and some of those people clicked through. Generative AI can use the information to become the destination itself. The legal battle now asks what happens when the source is indispensable to the answer but no longer indispensable to the visit.

0%