Guides7 min read

Where AI Labs Source Training Data in 2026

Troveo Team

Troveo

Every AI model is a bet on its training data, and in 2026 the question of where that data comes from has stopped being a technical detail and become a boardroom issue. Web archives that powered the last generation of models are largely tapped out, courts are picking apart how labs acquired their corpora, and the labs that keep shipping better models are the ones that keep finding data nobody else has.

Article banner reading Where Data Comes From, about how AI labs source training data

So where does training data actually come from? Six channels: scraped web data, open datasets, platform and user data, synthetic data, direct licensing deals, and licensed data providers. Nearly every lab uses a mix. What has changed is the mix itself.

The six sourcing channels

ChannelWhat it isBest forThe catch
Scraped web dataCommon Crawl and in-house crawlersBase pretraining corporaExhausted, commoditized, legally eroding
Open datasetsAcademic, government, and community corporaResearch, evals, fine-tuningEveryone has them; provenance often unclear
Platform and user dataData generated on platforms you own or rentConversation, behavioral, and video dataClosed to labs without a platform or a deal
Synthetic dataModel-generated training dataEdge cases, reasoning, scarce domainsInherits the ceiling of the generating model
Direct licensing dealsOne-to-one deals with major rights holdersBrand-name corpora at frontier-lab scaleDoes not scale to thousands of rights holders
Licensed data providersMarketplaces aggregating cleared rights holdersScarce, non-public, training-ready dataQuality varies; vet rights documentation
The six channels AI labs use to source training data in 2026

1. Scraped web data. Common Crawl and in-house crawlers built the foundation model era. The trouble is twofold. First, the open web is close to exhausted as a source of new high-quality text, and every lab already has it, so it moves no benchmarks. Second, the legal ground under scraping is eroding. The AI training data lawsuits keep turning on how data was acquired, and Anthropic's $1.5 billion settlement was driven by acquisition, not by training itself. Scraping still happens everywhere, but no serious lab treats it as a strategy anymore.

2. Open datasets. Academic corpora, government data, and open repositories like Hugging Face remain useful for research, evaluation, and fine-tuning. Their limits are the same as scraped data: everyone has them, and their provenance is often murkier than their licenses suggest, since many open datasets were themselves assembled by scraping.

3. Platform and user data. Labs attached to platforms train on what their users generate. Meta has its social graph, Google has YouTube, and companies without platforms pay for access to someone else's: Google pays Reddit roughly $60 million a year for its conversation data. This channel is powerful but closed. You either own a platform or you rent one.

4. Synthetic data. Model-generated data now fills real gaps, especially for edge cases, reasoning traces, and domains where real data is scarce or sensitive. But it inherits the ceiling of the models that generate it, and labs have learned to treat it as a supplement rather than a diet. Synthetic data works best layered on top of scarce real-world data, not instead of it.

5. Direct licensing deals. The headline channel of the last two years. News Corp licensed to OpenAI for a reported $250 million over five years. Amazon pays The New York Times a reported $20 to 25 million a year. Disney invested $1 billion in OpenAI alongside a deal allowing Sora to generate video with Disney characters. Music labels settled their lawsuits against Suno and Udio by converting them into licensing agreements. These deals work when you are a frontier lab and the rights holder is enormous. They do not scale down: negotiating one-off deals with thousands of smaller rights holders is a job no lab wants.

6. Licensed data providers and marketplaces. The channel built to solve exactly that problem. Instead of negotiating with thousands of fragmented rights holders, a lab works with one company that has already aggregated them, cleared the rights for AI training, and delivered the data cleaned and normalized. This is where Troveo operates, and the category matters most for the data that direct deals and scraping cannot reach: scarce, non-public video, audio, and gameplay data held by thousands of individual creators and media companies rather than a few big publishers.

What labs actually pay

The market has real price discovery now. Beyond the flagship deals above, the reported figures run from Axel Springer at around $13 million over three years, to the Financial Times at $5 to 10 million a year, to Wiley's academic licensing at over $40 million across two deals. Shutterstock generated $138 million in data licensing revenue in 2024 alone. Estimates put the AI training dataset market around $4 billion in 2026, with projections several times that by the early 2030s.

One shift inside the numbers is worth knowing: a growing share of publisher deals now cover retrieval and display rather than training, as publishers get more careful about permanent training rights. Deals for training-grade data increasingly happen quietly, and increasingly for modalities beyond text, where the open web never had much supply to begin with.

Why the mix is shifting

Three forces are pushing labs from free data toward licensed data.

The first is performance. Web text is commoditized, so the gains now come from data other labs do not have: multi-camera video, high-end animation, hours of continuous gameplay, professional audio, proprietary business data. Scarce beats big.

The second is legal. Courts keep separating how data was used from how it was acquired, and acquisition is where labs keep losing. A documented chain of rights has gone from nice-to-have to something legal teams require before a training run.

The third is the licensing market itself. Every new deal makes the market more real, and the more real the market, the weaker the fair use argument for taking data without a deal. The labs signing licensing deals are not just buying data. They are raising the legal cost of not buying it for everyone else.

How buyers choose a source

For teams evaluating where to get data, the practical questions are the same across channels: can the source prove chain of rights for every asset, do the licenses explicitly cover AI training, is the data exclusive or already in every competitor's corpus, and does it arrive training-ready or as a cleanup project. We cover that evaluation in detail in our guide to AI training data providers.

Where Troveo fits

Troveo is the licensed marketplace channel at scale: more than 7,000 licensors globally, 95 percent signed exclusively with Troveo, and over $20 million paid out to rights holders. Labs come to us for the data the other five channels cannot produce, real-world video, audio, gameplay, and business data that is rights-cleared for training, documented per asset, and delivered in the formats research teams actually use. You can browse and build datasets directly in Troveo Lens, or contact us and tell us what your team is training.

Frequently asked questions

Where do AI labs get their training data?
From six main channels: scraped web data, open datasets, platform and user data, synthetic data, direct licensing deals with rights holders, and licensed data providers or marketplaces. Most labs combine several, and the mix has been shifting toward licensed sources since the copyright lawsuits made data provenance a legal issue.
Do AI companies still scrape the web for training data?
Yes, but it no longer differentiates anyone. The open web is largely exhausted as a source of new high-quality training data, every lab already has it, and court rulings keep treating unauthorized acquisition as a separate legal problem from training. Scraping built the foundation model era; it is not where current gains come from.
How much do AI companies pay for training data?
Reported deals range from a few million a year for mid-size publishers to $250 million over five years for News Corp and OpenAI. Google pays Reddit roughly $60 million a year, Amazon pays The New York Times a reported $20 to 25 million a year, and Shutterstock earned $138 million from data licensing in 2024. Pricing for private training-data deals is usually confidential.
What is the difference between a data labeling company and a data licensing marketplace?
A labeling company like Scale AI processes and annotates data you already have or that it collects for you. A licensing marketplace supplies the underlying data itself, aggregated from rights holders who have explicitly cleared it for AI training. Many labs use both: one for the data, one for preparing it.
Why are AI labs buying licensed data instead of scraping it?
Three reasons. Scarce licensed data moves benchmarks when commodity web data no longer does. Courts keep ruling that how data was acquired matters legally, independent of fair use. And the growth of the licensing market itself weakens the fair use defense for unlicensed training, which raises the risk of relying on scraped data every year.
What kinds of licensed training data do labs buy?
Increasingly, data beyond text: real-world video for video generation and world models, professional audio, continuous gameplay footage, and proprietary business data. These modalities barely exist on the open web at training quality, so licensing is often the only way to acquire them at scale.
Is synthetic data replacing real training data?
No. Synthetic data fills gaps, especially for edge cases and reasoning tasks, but it is generated by models and inherits their limitations. Labs treat it as a layer on top of scarce real-world data. The industry pattern in 2026 is synthetic plus licensed, not synthetic instead of licensed.
How do AI labs find scarce or non-public data?
Mostly through licensed marketplaces and data partnerships, because scarce data is scarce precisely because it is not on the open web. It sits with thousands of individual creators, studios, and companies. Marketplaces aggregate those rights holders into one counterparty, clear the training rights, and deliver the data in training-ready formats.

Related articles

Back to Resources