Guides4 min read

How Much Does AI Training Data Cost in 2026?

Troveo Team

Troveo

There is no price list for AI training data, and anyone quoting a universal rate per hour of video or per thousand messages is guessing. Most contracts are confidential, pricing varies enormously with what the data is and who wants it, and the market is young enough that comparable deals are scarce. What does exist is a growing set of disclosed numbers that sketch the market's edges, and a consistent set of factors that move price inside them. This guide covers both.

Article banner reading What Data Costs, on AI training data pricing

The disclosed numbers

The public record runs from single-digit millions to a quarter billion, and the range itself is the lesson.

DealReported value
News Corp and OpenAI$250M+ over five years (2024)
Reddit and Google~$60M per year
Reddit and OpenAI~$70M per year (reported)
Shutterstock AI licensing revenue$104M in 2023, across five major buyers
Amazon and The New York Times$20-25M per year (reported)
OpenAI and Dotdash Meredith~$16M per year (reported)
Spirit Airlines enterprise data auction$10M (Google, 2026, pending approval)
Reported AI training data deal values, 2023 to 2026.

Two things stand out. First, famous catalogs command famous prices: the biggest numbers belong to brand-name publishers with massive, unique corpora. Second, even the "small" end is real money: a defunct airline's internal archive drew a $10 million bid, our breakdown of the Spirit Airlines data deal covers why. The full set of market figures, updated monthly, lives in our AI training data statistics page.

What actually moves the price

Underneath the headline deals, pricing follows a consistent logic. Scarcity is the biggest lever: data that exists nowhere else, real gameplay footage, first-person task video, specialized industry workflows, prices entirely differently from data with substitutes. Exclusivity multiplies it: a dataset your competitors can also license is worth a fraction of one they cannot, which is why exclusive catalogs command premiums. Rights cleanliness is priced in both directions: documented, per-asset, training-specific permission raises value, while murky provenance discounts a dataset toward worthless, since the legal risk transfers to the buyer, a dynamic our guide to rights-cleared training data explains. Readiness matters: cleaned, normalized, annotated, training-ready data is worth more than raw exports, because the buyer's alternative is doing that work themselves. And modality sets the baseline: scarce real-world video and audio price differently from text, which the open web made abundant.

Demand timing sits on top of all of it. Prices follow what labs are training right now: world models pulled up gameplay and first-person footage, voice agents pulled up natural conversation, and agentic AI is currently pulling up enterprise workflow data.

How licensing pricing is structured

Most licensed data transactions take one of a few shapes: flat licenses for a defined dataset and use, per-unit pricing by hour, asset, or volume tier, recurring annual licenses for refreshed or ongoing data, and revenue-share arrangements in some marketplaces where owners get paid as their data licenses. Terms matter as much as headline price: what uses are permitted, for which models, for how long, and whether the license is exclusive. A cheaper non-exclusive license and a pricier exclusive one are different products, not different prices for the same thing.

For buyers, the practical takeaway from our guide to how to buy AI training data applies here: name your gap precisely before you shop, because the price of "video data" is meaningless until you specify which video, with what rights, at what readiness.

What this means for data owners

The same factors run in reverse if you own data. Scarcity, exclusivity, clean rights, and readiness are what make your catalog worth licensing, and there is no universal price per message or per hour, which means serious valuation starts with an inventory of what you actually hold, not a quote. That process, and what owners can realistically expect, is covered in our guide to licensing company data for AI.

Where Troveo fits

Troveo sits in the middle of this market: a licensed data marketplace where more than 7,000 rights holders have licensed real-world video, audio, gaming, robotics, and business data, with more than 20 million dollars paid out to owners. For buyers, that aggregation is the alternative to negotiating headline deals one rights holder at a time: one agreement, documented rights per asset, training-ready delivery, at marketplace rather than News Corp prices. Browse the catalog in Lens or talk to us about what your budget gets in your modality.

Frequently asked questions

How much does AI training data cost?
There is no universal rate. Disclosed deals run from about $10 million for a company's archive to more than $250 million over five years for a major publisher's catalog, while marketplace licensing prices far below headline deals. Cost depends on scarcity, exclusivity, rights, readiness, and modality.
How much do AI companies pay for licensed content?
Reported figures include News Corp and OpenAI at more than $250 million over five years, Google paying Reddit around $60 million a year, Amazon and The New York Times at a reported $20 to 25 million a year, and Shutterstock earning $104 million from AI licensing in 2023.
Why is training data pricing so secretive?
Most contracts carry confidentiality terms, and both sides benefit: labs avoid setting precedents that raise the price of the next deal, and sellers avoid anchoring below what the next buyer might pay. The disclosed numbers mostly surface through securities filings, court records, and reporting.
What makes training data expensive?
Scarcity, exclusivity, clean and documented rights, and training readiness. Data that exists nowhere else, that competitors cannot also license, with per-asset permission and delivery-ready formatting, commands the premium. Abundant, non-exclusive, murky, or raw data discounts on every axis.
Is video training data more expensive than text?
Generally, real-world video and audio price above text, because the open web made text abundant while training-grade video and audio remain scarce. Within video, scarce categories like gameplay, first-person footage, and task demonstrations price above generic footage.
How is licensed data priced: per hour, per asset, or flat?
All three exist: flat licenses for defined datasets, per-unit pricing by hour or asset or volume tier, recurring licenses for ongoing data, and revenue shares in some marketplaces. Permitted uses, duration, and exclusivity move the effective price as much as the unit rate.
How much is my company's data worth?
It varies enormously, and anyone quoting a rate without seeing your data is guessing. Value depends on scale, history, uniqueness, cross-system context, visible outcomes, and rights cleanliness. Serious valuation starts with an inventory and a rights review, not a price list.
Is licensed data more expensive than scraping?
Per unit, yes. All-in, the comparison has shifted: courts keep finding liability in unlicensed acquisition, the Anthropic settlement alone cost $1.5 billion, and legal exposure, discovery costs, and dataset destruction orders are now part of scraping's real price.

Related articles

Back to Resources