A Court May Soon Decide Whether Training ChatGPT on Publisher Content Is Fair Use

A Court May Soon Decide Whether Training ChatGPT on Publisher Content Is Fair Use
Sponsored

One of the most important unresolved questions in generative AI may soon receive its clearest judicial answer yet: can a company copy copyrighted journalism and books without permission in order to train a large language model and still call that use fair?

OpenAI, Microsoft, The New York Times and a group of prominent authors have now put that question directly before a federal judge. Both sides filed competing motions for summary judgment on September 4, asking U.S. District Judge Sidney H. Stein in Manhattan to rule on OpenAI and Microsoft’s fair-use defense without waiting for a full trial on the issue.

The development was detailed by Reuters on September 8. The consolidated litigation brings together the Times’ long-running case over the use of its journalism and separate claims from authors including John Grisham, Jonathan Franzen and George R.R. Martin.

The stakes extend far beyond ChatGPT. If Judge Stein squarely decides whether this kind of model training qualifies as fair use, the ruling could become one of the most consequential early guideposts for how U.S. copyright law applies to the data underlying generative AI.

It would not necessarily end the national debate. But it could move that debate from competing legal theories toward an actual judicial rule applied to one of the industry’s most important AI systems.

Both sides are asking the judge to decide now

Summary judgment is a procedural mechanism that allows a court to resolve a legal claim or defense without a trial when the relevant material facts do not require a jury to settle them.

Here, the parties are effectively telling the court that the record developed through years of litigation is sufficient for Judge Stein to decide the fair-use issue in their favor.

OpenAI confirms on its official New York Times litigation page that it filed summary-judgment memoranda on September 4 addressing both the author and news-publisher disputes. The Authors Guild likewise says its class plaintiffs moved for summary judgment, asking the court to find infringement and reject OpenAI and Microsoft’s fair-use defense.

That does not mean a definitive ruling is guaranteed immediately. The judge can grant one side’s motion, grant only parts of a motion, deny both sides or conclude that factual disputes still require trial.

But the filings create a procedural path for the court to confront the central legal question directly.

The dispute begins with copying, not merely with what ChatGPT outputs

Public discussion of AI copyright disputes often focuses on whether a chatbot can reproduce a paragraph, imitate a writer or answer a question that might otherwise have sent a reader to a publisher’s website.

Those output questions matter, but the training dispute starts earlier.

The Times alleges that OpenAI and Microsoft used millions of copyrighted newspaper articles without authorization in the development of AI products including ChatGPT. The author plaintiffs similarly accuse the companies of using copyrighted books to train language models.

Training a large language model requires processing enormous amounts of source material. The legal question is whether making and using those copies for model training can be excused under the U.S. doctrine of fair use even when the copyright owner never granted a license.

That question is distinct from whether a particular ChatGPT response infringes a particular article or book.

A court could potentially view the training process as transformative while treating certain reproducing outputs differently. Conversely, market effects associated with AI outputs can become relevant to the fair-use analysis of training itself.

OpenAI says training is transformative analysis

OpenAI’s argument centers heavily on transformation.

In its public explanation of the case, the company says generative AI training is a transformative, non-expressive analytical use that does not substitute for the original works. Reuters reports that OpenAI told the court its pretraining process is designed to derive broad statistical patterns about language that can be used to create new text, rather than to reproduce protected expression.

That distinction is the foundation of the defense.

From OpenAI’s perspective, a newspaper article enters the training process as material from which a model learns patterns. The resulting model is not intended to function as a stored digital edition of that article. It is a general-purpose system capable of generating new outputs across an enormous range of tasks.

OpenAI therefore argues that the purpose of the copying is fundamentally different from the original purpose for which a journalist or author created the work.

The publishers say transformation does not erase market harm

The Times and other rights holders attack the problem from the opposite direction.

Their argument is not simply that copyrighted works were copied. They contend that the works were used commercially to build products that can compete with the very journalism and books on which the systems were trained.

Reuters reports that the news plaintiffs argue ChatGPT can divert users from publisher websites and displace markets for their work. The author plaintiffs similarly contend that generative AI can dilute the market for books and threaten the economic incentives that support human creation.

That moves the dispute toward another central component of fair use: the effect of the challenged use on the potential market for or value of the copyrighted work.

If AI training produces a genuinely different tool without substituting for the original, the defendants’ argument becomes stronger. If the resulting systems materially replace demand for the works, licensing markets or derivative uses, the plaintiffs’ argument becomes stronger.

The court may therefore have to decide not only what training does technically but what generative AI does economically.

Fair use is a four-factor test, not a rule that says “transformative equals legal”

U.S. copyright law evaluates fair use through four statutory factors.

Courts consider the purpose and character of the use, including its commercial nature; the nature of the copyrighted work; the amount and substantiality of the material used; and the effect of the use on the potential market for or value of the copyrighted work.

No single factor mechanically decides every case.

Transformative purpose has become especially influential in modern fair-use jurisprudence, but a court still considers the complete factual context. The Supreme Court’s copyright decisions have also cautioned against treating transformation as an unlimited license whenever a new product claims a different purpose.

That is why the OpenAI litigation matters. It asks a court to apply familiar copyright doctrine to a technological process that can ingest vast numbers of complete works to produce a general-purpose generative system.

Earlier AI training rulings point in OpenAI’s direction — but not without limits

The New York court is not writing on a completely blank slate.

Federal judges in California addressed related AI-training questions in 2025 in cases involving Anthropic and Meta. Reuters notes that Judge William Alsup described Anthropic’s use of books for training as “quintessentially transformative” in the circumstances before him.

Judge Vince Chhabria also found Meta’s training use transformative in a separate author lawsuit.

Those rulings gave AI developers significant support for the proposition that training a model can have a sufficiently different purpose from reading or distributing the original book.

But they did not create a blanket national rule that every form of AI training is fair use.

Chhabria, in particular, warned that generative AI could raise serious market concerns if it floods markets with works that compete with human-created material. The result of a fair-use case can depend on the evidence plaintiffs present about substitution and economic harm.

The New York litigation gives publishers and authors another opportunity to develop that record.

The New York Times case puts journalism’s business model directly into the analysis

The presence of major news publishers gives the dispute a dimension that is different from litigation focused only on books.

Digital journalism is sold through subscriptions, licensing, advertising and direct reader relationships. Publishers increasingly also negotiate licenses with AI companies for access to archives and real-time content.

If an AI company can obtain comparable material without paying and use it to build a commercial assistant that answers users’ questions, publishers argue that both audience and licensing markets can be affected.

That is why the dispute cannot be reduced to whether ChatGPT spits out verbatim copies of Times articles.

A publisher can argue that a sufficiently useful summary, synthesis or answer substitutes for the visit even when the language is not copied word for word.

OpenAI and Microsoft dispute that characterization and argue their products serve fundamentally different purposes rather than replacing journalism.

The judge will have to evaluate the evidence, not simply the rhetoric on either side.

Microsoft says the record does not show widespread substitution

Microsoft has emphasized evidence from the litigation suggesting that its AI products rarely reproduce substantial portions of the plaintiffs’ works.

As The Verge reported from Microsoft’s recent filings, the company points to analysis of millions of Copilot conversations that it says found very few examples containing substantial overlap with the books and news articles at issue.

Microsoft uses that evidence to argue that the products do not function as substitutes for copyrighted books or journalism.

The publishers reject that conclusion and say the discovery record supports their claim that OpenAI and Microsoft exploited copyrighted content to build competing commercial products.

The disagreement illustrates why output evidence matters even in a case centered on training. If the finished system frequently reproduces or replaces the original material, that can influence how a court understands the economic character of the underlying use.

The U.S. government has taken OpenAI’s side on fair use

The litigation also has an unusually significant government intervention.

The U.S. Justice Department filed a statement supporting OpenAI’s fair-use position. The government argued that training large language models on copyrighted material can provide substantial creative and public benefits and warned that overly restrictive copyright rules could impede AI development and U.S. competitiveness.

The filing adds institutional weight to OpenAI’s argument, but it does not decide the case.

The Justice Department is expressing the government’s legal position. Judge Stein remains responsible for applying copyright law to the evidentiary record before the court.

The Times and other rights holders strongly dispute the government’s view and argue that innovation policy cannot override the rights Congress granted to creators.

A ruling for OpenAI would strengthen the legal foundation of large-scale AI training

If the court holds that the training at issue is fair use, the decision would provide major support for AI companies that have built models using copyrighted material without obtaining a license for every work.

Such a ruling could reduce one of the largest legal uncertainties surrounding foundation-model development in the United States.

It would also strengthen the argument that the act of learning statistical patterns from copyrighted expression is legally distinct from distributing that expression.

But even a broad defense victory would not necessarily immunize every AI system from copyright liability.

Cases involving pirated source copies, memorized outputs, market substitution, different datasets or different product designs could still produce different results. Other claims in AI copyright litigation can also survive independently of a training fair-use defense.

Fair use remains fact-specific.

A ruling for publishers could transform the economics of AI development

The opposite result could be even more disruptive.

If the court concludes that unauthorized training on the plaintiffs’ works is not fair use, AI developers could face much greater pressure to license copyrighted material before using it in training datasets.

That could strengthen publishers, authors, record labels, image libraries and other rights holders seeking compensation from AI companies.

It could also increase the cost of building large models and favor companies capable of negotiating enormous licensing portfolios.

Existing models could face difficult questions about damages and remedies depending on the scope of any infringement finding.

The exact consequences would depend heavily on the wording of the ruling, the claims resolved and what happens on appeal.

The case could reshape the emerging AI licensing market

AI companies are already signing content deals with publishers and other rights holders even while arguing in court that training can qualify as fair use.

Those two positions are not necessarily inconsistent. A company can believe it has a legal right to train while still paying for higher-quality data, current information, product integration, attribution or reduced litigation risk.

But the bargaining power behind those negotiations depends partly on the legal baseline.

If unauthorized training is broadly protected, a publisher negotiates from one position. If a court says a license is required for the type of copying at issue, the publisher’s leverage changes dramatically.

The summary-judgment ruling could therefore influence contracts far beyond the parties in the courtroom.

Publishers are fighting over both inputs and outputs

The broader conflict between news organizations and generative AI has two connected fronts.

The first is the input: whether publishers should be compensated when their copyrighted archives are used to train models.

The second is the output: whether AI answers reduce the need for users to visit the publisher, subscribe or license the underlying work.

Those questions increasingly overlap.

A training process can be technologically transformative while the resulting product still competes economically with the source material. Conversely, a system can learn from a work without routinely reproducing it or eliminating the market for it.

Copyright law has to translate those realities into the four fair-use factors.

That is what makes this case more consequential than a dispute over whether a chatbot occasionally quotes too much text.

Summary judgment would be a major ruling, not the final word

Whatever Judge Stein decides, the phrase “a court decides whether AI training is fair use” needs qualification.

A federal district court ruling binds the parties before it and can be highly persuasive in other cases, especially when the factual record is extensive. It does not automatically become a nationwide rule equivalent to a Supreme Court decision.

The losing side can appeal. Other federal courts can confront different facts. Appellate courts may eventually disagree, creating a conflict that pushes the issue toward the Supreme Court.

The judge may also issue a narrower decision than either side wants.

For example, the court could resolve some fair-use factors while finding factual disputes on market harm, distinguish among datasets or plaintiffs, or leave other infringement theories for trial.

The procedural moment is therefore significant because a decisive ruling is possible, not because the entire copyright status of generative AI is guaranteed to be settled in one order.

The question is becoming too important for technology companies and publishers to answer for themselves

For several years, the AI industry and the creative industries have operated with fundamentally incompatible legal narratives.

AI developers say training is analogous to learning: models analyze works to identify patterns and use those patterns for new purposes. Publishers and authors say the analogy hides industrial-scale copying that extracts economic value from protected works without permission and then powers products capable of competing with their creators.

Both sides have enormous economic incentives to make their interpretation the legal default.

The New York litigation moves the dispute into a different phase because both camps are asking a judge to choose.

If Judge Stein reaches the merits of the fair-use defense on summary judgment, the resulting opinion could become one of the first major legal maps for the relationship between copyright and generative AI.

The ruling may favor OpenAI, the publishers, neither side completely, or only resolve part of the dispute. Appeals could follow for years.

But the central question is finally being presented in unusually direct form: when a company copies copyrighted journalism and books to teach a model like ChatGPT how language works, is that transformative fair use — or an unlicensed commercial exploitation of the works that made the system possible?

For AI companies and publishers alike, the answer could help determine who pays for the knowledge that trains the next generation of models.

0%