
The argument over artificial intelligence and copyright has moved from theory into courtrooms, government inquiries and billion-dollar business decisions.
In September 2026, OpenAI and Anthropic urged Australia to reconsider its refusal to create a broad copyright exception for AI training. At the same time, OpenAI and Microsoft are fighting the New York Times and a group of authors in a major U.S. case over whether training generative AI on copyrighted books and journalism can qualify as fair use.
That leaves one deceptively simple question:
Is it legal to train AI on copyrighted content?
The honest answer is: sometimes, depending on the country, how the material was obtained, how it was used, what the model produces and how the use affects the market for the original work.
There is no single global rule, and even in the United States the most important questions are still being tested in court.
This article explains the controversy in plain English, without assuming that either creators or AI companies have already won the argument.
Why this fight suddenly matters more
Modern generative AI systems are trained on enormous collections of text, images, audio, code and other material. Some of that material is public domain. Some is licensed. Some is user-generated. Some is synthetic. And some is protected by copyright.
That last category creates the conflict.
Creators argue that companies should not be able to ingest copyrighted work at enormous scale without permission, especially when the resulting systems can generate material that competes with the people whose work helped train them.
AI companies argue that training is fundamentally different from republishing a book, article or image. Their position is that a model analyzes patterns in data to build a new statistical system rather than storing a library for users to read.
The law has to decide how old copyright principles apply to a technical process that did not exist when many of those rules were written.
The short answer in the United States
In the United States, the central legal question is usually fair use.
Fair use can permit certain uses of copyrighted material without permission. Courts evaluate four broad factors:
- The purpose and character of the use.
- The nature of the copyrighted work.
- The amount and substantiality of the material used.
- The effect of the use on the market for the original work.
There is no automatic rule that says "AI training is fair use" or "AI training is infringement."
The U.S. Copyright Office's Part 3 report on generative AI training says the analysis is highly fact-specific. It recognizes that some training uses may be transformative, while also warning that commercial use of unlawfully sourced works or use that produces competing expressive content can weigh against fair use.
That is why apparently similar lawsuits can produce different outcomes.
Training is not the same legal question as output
One of the most confusing parts of the debate is that people often collapse several separate issues into one.
Consider these four stages:
- How the training data was acquired.
- Whether copying the data for training is legally permitted.
- What the trained model retains or memorizes.
- What the model later produces for users.
Those are related questions, but they are not identical.
A company might have a stronger argument that training is transformative while still facing a problem because the source material was pirated.
A model might be trained lawfully but later generate an output that is substantially similar to a protected work.
Or a model could produce original-looking output while the training process itself remains disputed.
Treating all of those situations as one question makes the legal debate look simpler than it really is.
Why the source of the training data matters
The source of the material can be crucial.
A court may analyze a lawfully purchased or publicly accessible copy differently from material obtained through an unauthorized shadow library or another illegal source.
The Anthropic copyright litigation illustrated this distinction. A court treated the training use itself differently from the company's acquisition of pirated books, and the later settlement focused heavily on those pirated source copies.
That distinction matters because "the model learned from a book" and "the company lawfully obtained the book" are not the same claim.
For developers, publishers and creators, this means data provenance is becoming a major compliance issue.
Knowing what was used, where it came from and under what rights is no longer just a research detail. It can be central evidence.
What OpenAI argues
OpenAI has publicly argued that training on publicly available internet material can qualify as fair use.
Its basic position is that model training is transformative because the system is not supposed to provide users with substitute copies of the original works. Instead, the model extracts patterns, relationships and statistical structure that can support many new uses.
OpenAI also points to economic and scientific benefits from AI and argues that a legal framework that allows training is important for innovation.
At the same time, OpenAI offers mechanisms for some content owners to opt out or express preferences about use of their material.
That does not settle the legal question. It is the company's position in an active policy and litigation debate.
What publishers and authors argue
Publishers, authors and other rights holders make a different argument.
They contend that large AI companies copied valuable creative works without permission and then built commercial systems that can compete with the people who created those works.
For many creators, the strongest concern is not only direct reproduction.
It is market substitution.
If a user can ask an AI system to summarize a paywalled article, imitate a style, draft a competing story or produce substitute content, creators argue that the economic value of the original work may be reduced.
The U.S. Copyright Office has also identified market effect as an important part of the analysis.
That is why the debate is not only about whether a model can technically reproduce a paragraph. It is also about whether the system changes the market in which the original creator earns money.
Why the New York Times case is so important
The consolidated New York copyright litigation involving the New York Times, authors, OpenAI and Microsoft is one of the clearest tests of these arguments.
The plaintiffs say copyrighted books and journalism were copied without permission and used to build systems that can compete with the same markets.
OpenAI and Microsoft argue that the training process is transformative and does not substitute for the originals in the way copyright law is designed to prevent.
In September 2026, both sides sought summary judgment on major parts of the dispute.
A court ruling could clarify how fair use applies to large-scale generative AI training, but it would not automatically settle every AI copyright case. Different datasets, acquisition methods, models and outputs can produce different legal facts.
For now, the safest description is that the law is still developing.
Why Australia is taking a different path
The controversy is not limited to the United States.
Australia has resisted creating a broad copyright exception that would let AI companies train on protected Australian content without permission.
OpenAI and Anthropic have argued that Australia should reconsider that position and create a more flexible framework for AI training.
Creator groups have pushed in the opposite direction, arguing that writers, artists, publishers and other rights holders should retain control over whether their work is used and whether compensation is required.
This disagreement shows why there is no single global answer to AI training.
A practice that an American company believes is protected by fair use may face a different legal framework in another country.
What about Europe and the UK?
European law uses a different structure from U.S. fair use.
The European Union has text-and-data-mining rules that can permit certain computational uses while also giving rights holders ways to reserve rights in some circumstances.
The United Kingdom has had a narrower exception for computational analysis for non-commercial research where lawful access exists. Proposals to broaden the framework for commercial AI have generated strong debate.
The practical lesson is simple:
AI developers cannot assume that a training practice allowed in one jurisdiction is automatically allowed everywhere.
Global systems need jurisdiction-aware data governance.
Is "publicly available" the same as "free to use"?
No.
This is one of the most common misunderstandings.
A webpage can be publicly accessible and still be protected by copyright.
The fact that anyone can read an article does not automatically mean anyone can reproduce it for any commercial purpose.
At the same time, public availability can be relevant to some legal arguments, including expectations, access and the character of a use.
The final legal analysis depends on more than whether a URL can be opened without a password.
Public access and copyright permission are different concepts.
Does robots.txt solve the copyright problem?
Not by itself.
Robots.txt is a technical instruction used by websites to communicate crawling preferences to automated systems.
It can be an important signal. Some AI companies also provide specific opt-out controls or crawler identities.
But copyright is a legal right, not just a crawling preference.
A crawler obeying robots.txt does not automatically prove that every use of the data is lawful. And a crawler ignoring robots.txt does not automatically answer every copyright question.
The two issues can overlap, but they are not interchangeable.
Can creators opt out of AI training?
Sometimes, but the effectiveness of an opt-out depends on the company, crawler, dataset and jurisdiction.
Some AI companies let website owners block specific crawlers. Some offer forms or policy controls. Publishers may also negotiate licensing agreements.
But an opt-out today may not remove copies already included in older datasets.
That creates a timing problem: the creator may be able to control future collection without fully controlling past ingestion.
For website owners, it is useful to keep records of crawler settings, licensing agreements and content ownership. Those records may matter later even if the law changes.
What counts as market harm?
Market harm is one of the hardest parts of the fair-use debate.
Suppose an AI model was trained on thousands of news articles but never reproduces one word-for-word.
Has the publisher lost anything?
The answer may depend on what the model does next.
If the AI simply learns general language patterns and helps a user write an unrelated email, the connection to the article's market may be weak.
If the AI can generate detailed substitutes for paywalled reporting, summarize exclusive journalism in a way that reduces visits, or flood a market with close substitutes, the market argument becomes stronger.
This is one reason publishers care so much about traffic.
AI is not only changing content production. It can also change how readers discover information.
That makes the copyright debate closely connected to the future of search, referrals and online publishing.
Why this matters to small bloggers too
It is easy to assume that AI copyright is a fight between giant technology companies and giant media companies.
Small publishers are affected too.
A blogger may use AI to research, outline, translate or edit an article. That does not automatically make the article infringing.
The bigger questions are:
- Did the writer copy protected material?
- Are sources properly attributed where needed?
- Is the final article genuinely original?
- Is the AI being asked to imitate a living creator too closely?
- Is third-party material being reproduced beyond what the law permits?
- Are images, screenshots and quotations used with appropriate rights or exceptions?
For a small site, the safest editorial strategy is not to obsess over whether a detector thinks text "looks AI."
The more important goals are originality, factual verification, useful analysis, clear sourcing and respect for other people's protected work.
That is also better for readers.
What AI users often misunderstand about copyright
There are several myths worth clearing up.
Myth 1: "If AI generated it, nobody can complain"
False.
AI-generated output can still create legal problems if it reproduces or is substantially similar to protected material.
Myth 2: "If it was online, it was free"
Also false.
Public availability does not erase copyright.
Myth 3: "Training and output are the same thing"
They are connected but legally distinct.
Myth 4: "Fair use means anything transformative is legal"
Not automatically.
Transformative purpose is important, but courts still weigh other factors, including market effect.
Myth 5: "One court ruling will settle all AI training"
Unlikely.
Different cases involve different datasets, models, jurisdictions and alleged harms.
Could licensing become the practical middle ground?
Licensing is already becoming part of the AI economy.
Some publishers and platforms have signed agreements that allow AI companies to use content under negotiated terms.
Licensing can give rights holders compensation and give AI developers clearer legal certainty.
But licensing also raises practical questions.
Who gets paid?
How is the value of millions of works measured?
Can a small creator negotiate on equal terms with a giant technology company?
What happens when the training dataset contains material from countries with different laws?
The U.S. Copyright Office has noted that voluntary licensing markets are developing but remain uneven.
That suggests the future may not be a simple choice between "everything is free to train on" and "nothing can be used."
Different sectors may end up with different licensing systems.
The piracy question may be easier than the fair-use question
One pattern is already becoming clearer.
Courts may be willing to consider whether lawful copies can be used for transformative training under fair use.
They have been much less sympathetic to illegal acquisition.
That makes piracy a separate risk.
An AI company cannot necessarily solve a sourcing problem by arguing that the later training step was transformative.
For anyone building datasets, "where did this file come from?" can matter as much as "what did the model do with it?"
That distinction is likely to remain important even as the broader fair-use law develops.
What should readers watch next?
Three developments matter most.
First, the New York federal litigation could produce a major ruling on whether specific forms of generative-AI training qualify as fair use.
Second, countries are continuing to choose different copyright frameworks for AI. Australia, the EU, the UK and the U.S. are not following identical paths.
Third, licensing markets may mature faster than legislation. Commercial deals between publishers and AI companies could shape practical behavior before courts answer every legal question.
For creators, that means today's uncertainty is real.
For AI companies, it means data governance is no longer only a technical problem.
For users, it means an AI answer may sit on top of a much larger legal and economic debate about where the model's knowledge came from.
How this connects to the wider AI-agent debate
Copyright is only one part of a larger shift in AI accountability.
Recent autonomous-agent incidents have raised questions about what happens when AI systems take actions beyond what their operators intended. We covered that separately in our explainer on whether AI agents can hack websites.
Shopping agents raise a different version of the same accountability question: who is responsible when an AI makes a bad decision with money or personal data? See our guide to AI shopping-agent fraud and privacy risks.
The common theme is that AI systems are moving from passive tools into systems that can influence markets, content, transactions and decisions.
Copyright is one of the first areas where that shift is being tested at scale.
FAQ
Is it definitely legal for ChatGPT or another AI model to train on copyrighted books?
No. In the United States, AI companies often rely on fair-use arguments, but whether a particular training use is lawful depends on the facts and is still being litigated. Other countries use different legal rules.
Has the New York Times already won its copyright case against OpenAI?
No. The major fair-use questions remain contested. The parties have asked the court to resolve key issues, but the litigation should not be described as a final victory for either side unless and until the court enters the relevant rulings.
Does buying a book give an AI company the right to train on it?
Buying or lawfully accessing a copy can matter, especially compared with using pirated copies, but it does not automatically resolve the separate fair-use question.
Can a website block AI crawlers?
Many website owners can use robots.txt, crawler-specific rules or provider opt-out tools. Those controls can affect future collection, but they do not by themselves answer all copyright questions or necessarily remove material already collected.
Is AI-generated content automatically copyright-free?
No. Copyright in AI-generated material depends on jurisdiction and the role of human authorship. Separately, an AI output can create infringement concerns if it reproduces protected expression.
Is using AI to help write a blog post illegal?
Not by itself. AI assistance can be used in legitimate writing workflows. The important questions include originality, human authorship where copyright protection is claimed, source verification and whether protected third-party material is copied beyond permitted limits.
Bottom line
The legal fight over AI training is not really one question.
It is a chain of questions:
Where did the data come from? Was it lawfully acquired? What exactly happened during training? What does the model output? Does that output substitute for the original market? And which country's law applies?
OpenAI and other AI developers argue that model training is transformative and often protected by fair use.
Publishers and creators argue that mass copying without permission can undermine the markets copyright law is meant to protect.
The U.S. Copyright Office has taken a more case-specific view: some uses may be transformative, but the answer can change depending on source, purpose, output controls and market effects.
That means the safest conclusion in 2026 is not that AI training is simply legal or illegal.
It is that the legal boundary is being drawn in real time.
This article is a general informational explainer, not legal advice.
Sources
- Reuters: Anthropic, OpenAI call for Australia to relax ban on training of AI models
- Reuters: OpenAI, New York Times case tees up key test of AI training under copyright law
- U.S. Copyright Office: Copyright and Artificial Intelligence, Part 3 — Generative AI Training
- OpenAI: OpenAI and journalism
- TechCrunch: Authors push back as publishers and agents make claims on Anthropic settlement