The training data question, unresolved
Do not expect an answer here
Whether training a model on copyrighted work without permission is lawful is being decided right now, in several countries, under different laws, with rulings pointing in different directions. This lesson gives you the shape of the disagreement so you can read the news about it. It does not resolve it, because it is not resolved, and any course that told you otherwise would be misleading you.
The claim on each side
Rights-holders argue that building a commercial product required copying their work; that the copying was at industrial scale; that some material was obtained from pirate sources; that models can reproduce protected expression; and that the output competes in the market for the originals, sometimes directly.
Developers argue that training extracts uncopyrightable statistical patterns rather than expression; that the use is transformative because the model does something different from the original work; that no copy is distributed to users; and that requiring permission for every work would make the technology impossible, which is an argument about consequences rather than doctrine.
Both positions contain something true, which is why this is hard.
What courts have actually said
A partial, deliberately non-conclusive list.
In Thomson Reuters v. Ross Intelligence (2025) a US court rejected a fair use defence where headnotes from a legal research service had been used to build a competing legal research tool. The tool was not generative, and the market-substitution argument was direct.
In Bartz v. Anthropic (2025) a US judge held that training a language model on books was transformative and qualified as fair use, while separately holding that downloading and retaining a library of pirated books was not — a split that separates the use from the acquisition. The case was subsequently reported to have settled for a very large sum in respect of the pirated-library claims.
In Kadrey v. Meta (2025) another US judge granted summary judgment to the developer on the record presented, while stating explicitly that the ruling did not establish that training on copyrighted work is generally lawful, and that a better-evidenced market-harm case might succeed.
In the UK, Getty Images' case against Stability AI proceeded to trial in 2025 with the principal training-related claims narrowed or dropped during proceedings, and the remaining judgment turned largely on trade mark and secondary infringement points rather than settling the training question.
In India, ANI's suit against OpenAI in the Delhi High Court, filed in 2024, is being watched closely, in part because India's fair dealing provisions are enumerated rather than open-ended, so a transformative-use argument of the American kind has less room to operate.
Read those together and the pattern is: courts are distinguishing acquisition from use, weighing market substitution heavily, and declining to announce general rules.
The legislative side
Japan has the most permissive position: Article 30-4 of its Copyright Act broadly permits use of works for information analysis where enjoyment of the expression is not the purpose, subject to limits.
The EU permits text and data mining under the 2019 Directive, with a research exception and a commercial exception that rights-holders may opt out of in a machine-readable way. The AI Act then requires general-purpose model providers to have a policy respecting those opt-outs and to publish a sufficiently detailed summary of training content. In practice the opt-out mechanism is contested: there is no single standard for expressing it and no easy way for a rights-holder to verify compliance.
The UK consulted in 2024–25 on a similar exception with a rights reservation, met sustained opposition from the creative industries, and the transparency question was fought over in Parliament during the passage of data legislation in 2025. The outcome remains unsettled.
The US has no legislation on this; it is being decided case by case.
What is reasonably safe to say
Four things, stated at the level of confidence they deserve.
Acquisition matters separately from use. Where copies were obtained unlawfully, that has been treated as a distinct wrong even by judges sympathetic to training as fair use. This is the clearest signal from the case law so far.
Market substitution is the decisive factor. Where the output competes directly with the input — a tool that replaces the licensed product it learned from — defences have fared worse.
Verbatim reproduction is a separate problem. Memorised output is expression, not pattern, and every framework treats it differently from training.
Your jurisdiction may reach a different answer, and if you are relying on this in a business, that is a question for a lawyer in your country rather than an inference from an American headline.
What to do in the meantime
If you build on generated material commercially: prefer providers offering indemnities and read what they cover; keep records of what was generated when and with what; avoid prompting for named living artists or protected characters; and treat any output that looks like something you recognise as a warning rather than a success.
The one thing to keep
Courts are separating unlawful acquisition from transformative use and weighing market substitution heavily, legislatures are diverging from Japan's broad permission to India's enumerated fair dealing, and no jurisdiction has settled the general question.
Before you move on
A US ruling found that training a language model on books was transformative fair use, while separately finding that downloading a pirated library was not. What is the significance of that split?
Pick the one you would defend. Nobody sees your answer.