There is no price list for AI training data, and anyone quoting a universal rate per hour of video or per thousand messages is guessing. Most contracts are confidential, pricing varies enormously with what the data is and who wants it, and the market is young enough that comparable deals are scarce. What does exist is a growing set of disclosed numbers that sketch the market's edges, and a consistent set of factors that move price inside them. This guide covers both.
The disclosed numbers
The public record runs from single-digit millions to a quarter billion, and the range itself is the lesson.
| Deal | Reported value |
|---|---|
| News Corp and OpenAI | $250M+ over five years (2024) |
| Reddit and Google | ~$60M per year |
| Reddit and OpenAI | ~$70M per year (reported) |
| Shutterstock AI licensing revenue | $104M in 2023, across five major buyers |
| Amazon and The New York Times | $20-25M per year (reported) |
| OpenAI and Dotdash Meredith | ~$16M per year (reported) |
| Spirit Airlines enterprise data auction | $10M (Google, 2026, pending approval) |
Two things stand out. First, famous catalogs command famous prices: the biggest numbers belong to brand-name publishers with massive, unique corpora. Second, even the "small" end is real money: a defunct airline's internal archive drew a $10 million bid, our breakdown of the Spirit Airlines data deal covers why. The full set of market figures, updated monthly, lives in our AI training data statistics page.
What actually moves the price
Underneath the headline deals, pricing follows a consistent logic. Scarcity is the biggest lever: data that exists nowhere else, real gameplay footage, first-person task video, specialized industry workflows, prices entirely differently from data with substitutes. Exclusivity multiplies it: a dataset your competitors can also license is worth a fraction of one they cannot, which is why exclusive catalogs command premiums. Rights cleanliness is priced in both directions: documented, per-asset, training-specific permission raises value, while murky provenance discounts a dataset toward worthless, since the legal risk transfers to the buyer, a dynamic our guide to rights-cleared training data explains. Readiness matters: cleaned, normalized, annotated, training-ready data is worth more than raw exports, because the buyer's alternative is doing that work themselves. And modality sets the baseline: scarce real-world video and audio price differently from text, which the open web made abundant.
Demand timing sits on top of all of it. Prices follow what labs are training right now: world models pulled up gameplay and first-person footage, voice agents pulled up natural conversation, and agentic AI is currently pulling up enterprise workflow data.
How licensing pricing is structured
Most licensed data transactions take one of a few shapes: flat licenses for a defined dataset and use, per-unit pricing by hour, asset, or volume tier, recurring annual licenses for refreshed or ongoing data, and revenue-share arrangements in some marketplaces where owners get paid as their data licenses. Terms matter as much as headline price: what uses are permitted, for which models, for how long, and whether the license is exclusive. A cheaper non-exclusive license and a pricier exclusive one are different products, not different prices for the same thing.
For buyers, the practical takeaway from our guide to how to buy AI training data applies here: name your gap precisely before you shop, because the price of "video data" is meaningless until you specify which video, with what rights, at what readiness.
What this means for data owners
The same factors run in reverse if you own data. Scarcity, exclusivity, clean rights, and readiness are what make your catalog worth licensing, and there is no universal price per message or per hour, which means serious valuation starts with an inventory of what you actually hold, not a quote. That process, and what owners can realistically expect, is covered in our guide to licensing company data for AI.
Where Troveo fits
Troveo sits in the middle of this market: a licensed data marketplace where more than 7,000 rights holders have licensed real-world video, audio, gaming, robotics, and business data, with more than 20 million dollars paid out to owners. For buyers, that aggregation is the alternative to negotiating headline deals one rights holder at a time: one agreement, documented rights per asset, training-ready delivery, at marketplace rather than News Corp prices. Browse the catalog in Lens or talk to us about what your budget gets in your modality.
