Every AI model is a bet on its training data, and in 2026 the question of where that data comes from has stopped being a technical detail and become a boardroom issue. Web archives that powered the last generation of models are largely tapped out, courts are picking apart how labs acquired their corpora, and the labs that keep shipping better models are the ones that keep finding data nobody else has.
So where does training data actually come from? Six channels: scraped web data, open datasets, platform and user data, synthetic data, direct licensing deals, and licensed data providers. Nearly every lab uses a mix. What has changed is the mix itself.
The six sourcing channels
| Channel | What it is | Best for | The catch |
|---|---|---|---|
| Scraped web data | Common Crawl and in-house crawlers | Base pretraining corpora | Exhausted, commoditized, legally eroding |
| Open datasets | Academic, government, and community corpora | Research, evals, fine-tuning | Everyone has them; provenance often unclear |
| Platform and user data | Data generated on platforms you own or rent | Conversation, behavioral, and video data | Closed to labs without a platform or a deal |
| Synthetic data | Model-generated training data | Edge cases, reasoning, scarce domains | Inherits the ceiling of the generating model |
| Direct licensing deals | One-to-one deals with major rights holders | Brand-name corpora at frontier-lab scale | Does not scale to thousands of rights holders |
| Licensed data providers | Marketplaces aggregating cleared rights holders | Scarce, non-public, training-ready data | Quality varies; vet rights documentation |
1. Scraped web data. Common Crawl and in-house crawlers built the foundation model era. The trouble is twofold. First, the open web is close to exhausted as a source of new high-quality text, and every lab already has it, so it moves no benchmarks. Second, the legal ground under scraping is eroding. The AI training data lawsuits keep turning on how data was acquired, and Anthropic's $1.5 billion settlement was driven by acquisition, not by training itself. Scraping still happens everywhere, but no serious lab treats it as a strategy anymore.
2. Open datasets. Academic corpora, government data, and open repositories like Hugging Face remain useful for research, evaluation, and fine-tuning. Their limits are the same as scraped data: everyone has them, and their provenance is often murkier than their licenses suggest, since many open datasets were themselves assembled by scraping.
3. Platform and user data. Labs attached to platforms train on what their users generate. Meta has its social graph, Google has YouTube, and companies without platforms pay for access to someone else's: Google pays Reddit roughly $60 million a year for its conversation data. This channel is powerful but closed. You either own a platform or you rent one.
4. Synthetic data. Model-generated data now fills real gaps, especially for edge cases, reasoning traces, and domains where real data is scarce or sensitive. But it inherits the ceiling of the models that generate it, and labs have learned to treat it as a supplement rather than a diet. Synthetic data works best layered on top of scarce real-world data, not instead of it.
5. Direct licensing deals. The headline channel of the last two years. News Corp licensed to OpenAI for a reported $250 million over five years. Amazon pays The New York Times a reported $20 to 25 million a year. Disney invested $1 billion in OpenAI alongside a deal allowing Sora to generate video with Disney characters. Music labels settled their lawsuits against Suno and Udio by converting them into licensing agreements. These deals work when you are a frontier lab and the rights holder is enormous. They do not scale down: negotiating one-off deals with thousands of smaller rights holders is a job no lab wants.
6. Licensed data providers and marketplaces. The channel built to solve exactly that problem. Instead of negotiating with thousands of fragmented rights holders, a lab works with one company that has already aggregated them, cleared the rights for AI training, and delivered the data cleaned and normalized. This is where Troveo operates, and the category matters most for the data that direct deals and scraping cannot reach: scarce, non-public video, audio, and gameplay data held by thousands of individual creators and media companies rather than a few big publishers.
What labs actually pay
The market has real price discovery now. Beyond the flagship deals above, the reported figures run from Axel Springer at around $13 million over three years, to the Financial Times at $5 to 10 million a year, to Wiley's academic licensing at over $40 million across two deals. Shutterstock generated $138 million in data licensing revenue in 2024 alone. Estimates put the AI training dataset market around $4 billion in 2026, with projections several times that by the early 2030s.
One shift inside the numbers is worth knowing: a growing share of publisher deals now cover retrieval and display rather than training, as publishers get more careful about permanent training rights. Deals for training-grade data increasingly happen quietly, and increasingly for modalities beyond text, where the open web never had much supply to begin with.
Why the mix is shifting
Three forces are pushing labs from free data toward licensed data.
The first is performance. Web text is commoditized, so the gains now come from data other labs do not have: multi-camera video, high-end animation, hours of continuous gameplay, professional audio, proprietary business data. Scarce beats big.
The second is legal. Courts keep separating how data was used from how it was acquired, and acquisition is where labs keep losing. A documented chain of rights has gone from nice-to-have to something legal teams require before a training run.
The third is the licensing market itself. Every new deal makes the market more real, and the more real the market, the weaker the fair use argument for taking data without a deal. The labs signing licensing deals are not just buying data. They are raising the legal cost of not buying it for everyone else.
How buyers choose a source
For teams evaluating where to get data, the practical questions are the same across channels: can the source prove chain of rights for every asset, do the licenses explicitly cover AI training, is the data exclusive or already in every competitor's corpus, and does it arrive training-ready or as a cleanup project. We cover that evaluation in detail in our guide to AI training data providers.
Where Troveo fits
Troveo is the licensed marketplace channel at scale: more than 7,000 licensors globally, 95 percent signed exclusively with Troveo, and over $20 million paid out to rights holders. Labs come to us for the data the other five channels cannot produce, real-world video, audio, gameplay, and business data that is rights-cleared for training, documented per asset, and delivered in the formats research teams actually use. You can browse and build datasets directly in Troveo Lens, or contact us and tell us what your team is training.
