If you work at an AI lab, the copyright lawsuits are no longer background noise. In July 2026, a federal judge gave final approval to Anthropic's $1.5 billion settlement with authors, the largest copyright payout in history. The same month, Hachette, Elsevier, and Cengage sued Google over books allegedly used to train Gemini. Between those two headlines sit dozens of active cases against nearly every major model developer.
Here is the part that matters if you buy training data: the labs that lost money did not lose because training on copyrighted work was ruled illegal. Courts have mostly said the opposite. They lost, or are still exposed, because of where the data came from. Provenance, not training, is what the lawsuits actually turn on.
This piece walks through where the major cases stand, what the rulings did and did not settle, and what that means for how you source data.
Is it legal to train AI on copyrighted data?
The unsatisfying but accurate answer: it depends on how you got the data, and courts do not agree with each other yet.
Two federal judges in California ruled in mid-2025 that training a large language model on copyrighted books can qualify as fair use. In Bartz v. Anthropic, Judge William Alsup called training "transformative, spectacularly so." In Kadrey v. Meta, the court also sided with the AI developer, though the judge went out of his way to say the authors lost because their market harm arguments were underdeveloped, not because the theory could never work.
But a third court went the other way. In Thomson Reuters v. Ross Intelligence, the court found that copying Westlaw headnotes to train a competing legal research tool was not fair use, largely because the output competed directly with the source. That case is now on appeal at the Third Circuit.
So the current state of the law is a split. Training can be fair use. It can also not be. The outcome depends on the facts, the court, and increasingly on one question: did you acquire the data legitimately?
The cases defining the landscape
| Case | Defendant | Status (July 2026) | What it turned on |
|---|---|---|---|
| Bartz v. Anthropic | Anthropic | Settled, $1.5B, final approval July 2026 | Training was fair use, pirated acquisition was not |
| Kadrey v. Meta | Meta | Fair use win; torrenting claims still active | Market harm arguments failed, data handling claims survived |
| In re OpenAI Copyright Litigation | OpenAI | Ongoing, 12 consolidated cases in SDNY | Discovery into outputs and data practices |
| Thomson Reuters v. Ross | Ross Intelligence | Ross lost; on appeal at Third Circuit | Output competed directly with the source data |
| Getty Images v. Stability AI | Stability AI | UK case largely failed; US case ongoing | Jurisdiction, training happened outside the UK |
| Disney v. Midjourney | Midjourney | Pending | Character outputs; up to $150K per work sought |
| Labels v. Suno and Udio | Suno, Udio | Settled into licensing deals, late 2025 | Unlicensed training converted to paid licenses |
| Publishers v. Google | Filed July 2026 | Books provided for search allegedly used for Gemini training |
Bartz v. Anthropic (settled, $1.5 billion). This is the case every data buyer should understand. Anthropic won the fair use argument on training. It still paid $1.5 billion, roughly $3,000 per work across about 500,000 books, because it had downloaded much of its library from pirate sites like Library Genesis. The court treated acquisition as a separate act from training. Fair use covered the training. Nothing covered the piracy. Final approval came in July 2026, and the settlement also required Anthropic to destroy the pirated datasets.
Kadrey v. Meta (partially resolved, partially alive). Meta won on fair use against the group of authors who sued, but claims over its torrenting activity, specifically "seeding" pirated files back to other users, are still active in 2026. Again the exposure that survived is about how the data was obtained and handled, not the training itself.
In re OpenAI Copyright Litigation (ongoing). Twelve consolidated cases, including The New York Times suit, are moving through discovery in the Southern District of New York. Courts have ordered OpenAI to produce tens of millions of output logs. Whatever the outcome, it has already shown that a lab's data practices will be examined in granular detail once litigation starts.
Thomson Reuters v. Ross (on appeal). The clearest loss for an AI developer so far, and a warning for anyone training models whose outputs compete with the data source.
Getty Images v. Stability AI. Getty's UK case largely failed in late 2025, mostly on jurisdictional grounds because the training happened outside the UK. Its US case continues.
Disney v. Midjourney (pending). Disney and Universal are seeking up to $150,000 per work for willful infringement over character images. Meanwhile Disney signed a licensing deal with OpenAI that lets Sora generate video with Disney characters. Suing one lab while licensing to another is not a contradiction. It is the market forming.
The music cases (settled into licensing deals). Universal and Warner sued Suno and Udio for training on recordings without permission, then settled in late 2025 by converting the disputes into licensing agreements. The lawsuits did not end AI music. They ended unlicensed AI music.
Publishers v. Google (filed July 2026). The newest major case alleges Google trained Gemini on books that publishers had provided to Google Books for search snippets only. The complaint cites an internal Google document warning that training on copyrighted books could mean "$10Bs-$100Bs in potential fines." Discovery will test that.
What the rulings actually turn on
Read the cases together and a pattern shows up that most coverage misses.
Courts keep separating two questions: whether training on a work is fair use, and whether the copy you trained on was lawfully acquired. A lab can win the first and still lose the second. That is exactly what happened to Anthropic, and it is the theory keeping the Meta torrenting claims alive. Fair use is a defense to how you use a work. It is not a defense to how you got it.
The second pattern is about market harm. Fair use analysis weighs whether the copying damages the market for the original work, and that includes the market for licensing it. In 2023 there was barely a licensing market for training data, which made harm hard to prove. In 2026 there is one. Disney licenses to OpenAI. Warner and Universal license to Suno and Udio. Every major lab has signed data licensing deals. Which creates an uncomfortable feedback loop for anyone relying on fair use: each new licensing deal makes the licensing market more real, and the more real that market is, the weaker the fair use defense becomes for training on unlicensed data. The legal ground under scraped data is not holding steady. It is eroding.
The third pattern is practical. Even the labs that won spent years in discovery, produced internal documents and output logs by the millions, and absorbed legal costs and headlines. For a frontier lab, the cost of defending a data provenance case now rivals the cost of just licensing the data.
What this means if you're buying training data
None of this says copyrighted material is off limits for training. It says unaccounted-for material is a liability. If you are evaluating data sources, the lawsuits translate into a short list of questions:
Where did each asset come from? Not the category, the chain. Who created it, who holds the rights, and what did they agree to? "Publicly available" is not a provenance answer, and the Google case shows that even data acquired legitimately for one purpose can create exposure when used for another.
Do the licenses actually cover AI training? Older content licenses often cover distribution or display but say nothing about model training. A rights-cleared dataset means the rights holder explicitly agreed to training use.
Can the provider document it? If a dispute or a due-diligence review ever happens, you want a paper trail per asset, not assurances.
Is the data exclusive or resold everywhere? Beyond the legal question there is a model quality question. Data that every lab already has moves no benchmarks.
This is the same evaluation logic covered in our guide to AI training data providers, and it applies whether you are sourcing video, audio, text, or gameplay data for world models.
Where Troveo fits
Troveo exists on the clean side of this landscape. We license video, audio, gameplay, and business data directly from the people who own it: more than 7,000 licensors globally, 95 percent of them signed exclusively with Troveo, with over $20 million paid out to rights holders so far. Every asset comes with a documented chain of rights that explicitly covers AI training, and the data is cleaned and normalized so teams get training-ready datasets instead of a legal review project.
The lawsuits are pushing the industry toward exactly this model, one deal at a time. Buying licensed data is no longer just the defensible choice. Increasingly it is the standard one.
If you want to see what licensed, rights-cleared data looks like in practice, you can browse and build datasets directly in Troveo Lens or contact us to talk through what your team is training.
