Licensing & Copyright9 min read

AI Training Data Lawsuits: What Buyers Need to Know in 2026

Troveo Team

Troveo

If you work at an AI lab, the copyright lawsuits are no longer background noise. In July 2026, a federal judge gave final approval to Anthropic's $1.5 billion settlement with authors, the largest copyright payout in history. The same month, Hachette, Elsevier, and Cengage sued Google over books allegedly used to train Gemini. Between those two headlines sit dozens of active cases against nearly every major model developer.

Article banner reading Data Is On Trial, about the AI training data copyright lawsuits

Here is the part that matters if you buy training data: the labs that lost money did not lose because training on copyrighted work was ruled illegal. Courts have mostly said the opposite. They lost, or are still exposed, because of where the data came from. Provenance, not training, is what the lawsuits actually turn on.

This piece walks through where the major cases stand, what the rulings did and did not settle, and what that means for how you source data.

Is it legal to train AI on copyrighted data?

The unsatisfying but accurate answer: it depends on how you got the data, and courts do not agree with each other yet.

Two federal judges in California ruled in mid-2025 that training a large language model on copyrighted books can qualify as fair use. In Bartz v. Anthropic, Judge William Alsup called training "transformative, spectacularly so." In Kadrey v. Meta, the court also sided with the AI developer, though the judge went out of his way to say the authors lost because their market harm arguments were underdeveloped, not because the theory could never work.

But a third court went the other way. In Thomson Reuters v. Ross Intelligence, the court found that copying Westlaw headnotes to train a competing legal research tool was not fair use, largely because the output competed directly with the source. That case is now on appeal at the Third Circuit.

So the current state of the law is a split. Training can be fair use. It can also not be. The outcome depends on the facts, the court, and increasingly on one question: did you acquire the data legitimately?

The cases defining the landscape

CaseDefendantStatus (July 2026)What it turned on
Bartz v. AnthropicAnthropicSettled, $1.5B, final approval July 2026Training was fair use, pirated acquisition was not
Kadrey v. MetaMetaFair use win; torrenting claims still activeMarket harm arguments failed, data handling claims survived
In re OpenAI Copyright LitigationOpenAIOngoing, 12 consolidated cases in SDNYDiscovery into outputs and data practices
Thomson Reuters v. RossRoss IntelligenceRoss lost; on appeal at Third CircuitOutput competed directly with the source data
Getty Images v. Stability AIStability AIUK case largely failed; US case ongoingJurisdiction, training happened outside the UK
Disney v. MidjourneyMidjourneyPendingCharacter outputs; up to $150K per work sought
Labels v. Suno and UdioSuno, UdioSettled into licensing deals, late 2025Unlicensed training converted to paid licenses
Publishers v. GoogleGoogleFiled July 2026Books provided for search allegedly used for Gemini training
Status of the major AI training data copyright cases as of July 2026

Bartz v. Anthropic (settled, $1.5 billion). This is the case every data buyer should understand. Anthropic won the fair use argument on training. It still paid $1.5 billion, roughly $3,000 per work across about 500,000 books, because it had downloaded much of its library from pirate sites like Library Genesis. The court treated acquisition as a separate act from training. Fair use covered the training. Nothing covered the piracy. Final approval came in July 2026, and the settlement also required Anthropic to destroy the pirated datasets.

Kadrey v. Meta (partially resolved, partially alive). Meta won on fair use against the group of authors who sued, but claims over its torrenting activity, specifically "seeding" pirated files back to other users, are still active in 2026. Again the exposure that survived is about how the data was obtained and handled, not the training itself.

In re OpenAI Copyright Litigation (ongoing). Twelve consolidated cases, including The New York Times suit, are moving through discovery in the Southern District of New York. Courts have ordered OpenAI to produce tens of millions of output logs. Whatever the outcome, it has already shown that a lab's data practices will be examined in granular detail once litigation starts.

Thomson Reuters v. Ross (on appeal). The clearest loss for an AI developer so far, and a warning for anyone training models whose outputs compete with the data source.

Getty Images v. Stability AI. Getty's UK case largely failed in late 2025, mostly on jurisdictional grounds because the training happened outside the UK. Its US case continues.

Disney v. Midjourney (pending). Disney and Universal are seeking up to $150,000 per work for willful infringement over character images. Meanwhile Disney signed a licensing deal with OpenAI that lets Sora generate video with Disney characters. Suing one lab while licensing to another is not a contradiction. It is the market forming.

The music cases (settled into licensing deals). Universal and Warner sued Suno and Udio for training on recordings without permission, then settled in late 2025 by converting the disputes into licensing agreements. The lawsuits did not end AI music. They ended unlicensed AI music.

Publishers v. Google (filed July 2026). The newest major case alleges Google trained Gemini on books that publishers had provided to Google Books for search snippets only. The complaint cites an internal Google document warning that training on copyrighted books could mean "$10Bs-$100Bs in potential fines." Discovery will test that.

What the rulings actually turn on

Read the cases together and a pattern shows up that most coverage misses.

Courts keep separating two questions: whether training on a work is fair use, and whether the copy you trained on was lawfully acquired. A lab can win the first and still lose the second. That is exactly what happened to Anthropic, and it is the theory keeping the Meta torrenting claims alive. Fair use is a defense to how you use a work. It is not a defense to how you got it.

The second pattern is about market harm. Fair use analysis weighs whether the copying damages the market for the original work, and that includes the market for licensing it. In 2023 there was barely a licensing market for training data, which made harm hard to prove. In 2026 there is one. Disney licenses to OpenAI. Warner and Universal license to Suno and Udio. Every major lab has signed data licensing deals. Which creates an uncomfortable feedback loop for anyone relying on fair use: each new licensing deal makes the licensing market more real, and the more real that market is, the weaker the fair use defense becomes for training on unlicensed data. The legal ground under scraped data is not holding steady. It is eroding.

The third pattern is practical. Even the labs that won spent years in discovery, produced internal documents and output logs by the millions, and absorbed legal costs and headlines. For a frontier lab, the cost of defending a data provenance case now rivals the cost of just licensing the data.

What this means if you're buying training data

None of this says copyrighted material is off limits for training. It says unaccounted-for material is a liability. If you are evaluating data sources, the lawsuits translate into a short list of questions:

Where did each asset come from? Not the category, the chain. Who created it, who holds the rights, and what did they agree to? "Publicly available" is not a provenance answer, and the Google case shows that even data acquired legitimately for one purpose can create exposure when used for another.

Do the licenses actually cover AI training? Older content licenses often cover distribution or display but say nothing about model training. A rights-cleared dataset means the rights holder explicitly agreed to training use.

Can the provider document it? If a dispute or a due-diligence review ever happens, you want a paper trail per asset, not assurances.

Is the data exclusive or resold everywhere? Beyond the legal question there is a model quality question. Data that every lab already has moves no benchmarks.

This is the same evaluation logic covered in our guide to AI training data providers, and it applies whether you are sourcing video, audio, text, or gameplay data for world models.

Where Troveo fits

Troveo exists on the clean side of this landscape. We license video, audio, gameplay, and business data directly from the people who own it: more than 7,000 licensors globally, 95 percent of them signed exclusively with Troveo, with over $20 million paid out to rights holders so far. Every asset comes with a documented chain of rights that explicitly covers AI training, and the data is cleaned and normalized so teams get training-ready datasets instead of a legal review project.

The lawsuits are pushing the industry toward exactly this model, one deal at a time. Buying licensed data is no longer just the defensible choice. Increasingly it is the standard one.

If you want to see what licensed, rights-cleared data looks like in practice, you can browse and build datasets directly in Troveo Lens or contact us to talk through what your team is training.

Frequently asked questions

Is it legal to train AI models on copyrighted data?
Sometimes. Two federal courts ruled in 2025 that training on copyrighted books can be fair use, while another found training was not fair use when the output competed with the source. All of the decisive rulings so far have hinged on whether the data was lawfully acquired. Training on licensed data avoids the question entirely.
What was the Anthropic copyright settlement?
Anthropic agreed to pay $1.5 billion to authors and publishers, about $3,000 per work across roughly 500,000 books, after a court found that downloading books from pirate sites violated copyright even though the training itself was fair use. A federal judge gave the settlement final approval in July 2026. It is the largest copyright settlement on record.
Did courts rule that AI training is fair use?
Two district courts said yes for LLM training on books, one court said no for a competing legal research tool, and appeals are pending. There is no nationwide rule yet, and the growth of the data licensing market is expected to make fair use harder to argue over time.
What is rights-cleared training data?
Data where the rights holder has explicitly licensed the content for AI training, with a documented chain of rights for every asset. It is different from scraped data, where rights status is unknown, and from generally licensed content, where the license may not cover training use.
Why does provenance matter so much in these cases?
Because courts have treated acquisition and training as separate acts. Anthropic won on training and still paid $1.5 billion over acquisition. Buyers inherit this logic: a model is only as defensible as the sourcing of its training data.
Does buying licensed data fully protect an AI company from lawsuits?
No sourcing strategy removes all litigation risk. But licensed data with documented, training-specific rights removes the acquisition and provenance issues that have driven the largest payouts so far.
How should I vet a training data provider on copyright?
Ask where each asset comes from, whether the license explicitly covers AI training, whether rights documentation exists per asset, and whether the provider actually pays rights holders. A provider who cannot answer those four questions quickly is a provider whose risk you are absorbing.
Do these lawsuits affect video and audio training data too?
Yes. The rulings so far mostly involve books and text, but the legal logic applies to any modality. Getty's cases target image training, Disney's suit targets video outputs, and the music settlements covered audio. Video and audio may carry more risk in practice, since a single clip can stack multiple rights: the footage itself, music, performers, and visible brands. Licensed sourcing resolves all of those at once.

Related articles

Back to Resources