A data licensing business model is one where a company earns revenue by granting others the right to use data it already generates, rather than by selling a product built on that data. In 2026 the buyers driving it are AI developers, who pay for proprietary, rights-cleared data because the public web no longer contains what their models need. Publishers, forums, stock libraries, an airline in bankruptcy, and a growing number of startups winding down have all run some version of the model. This guide explains what it is, the four ways it pays, why a business would choose it, what it takes to run, and when it is the wrong idea.
What the model actually is
Two things separate a data licensing model from simply selling a database. First, it is almost always licensing, not selling. The company keeps ownership and grants a defined right: to use defined data, for defined purposes such as training or evaluation, for a defined period, exclusively or not. The same data can be licensed more than once. Second, the data is a byproduct. It was produced while doing something else, running a newsroom, hosting a community, operating an airline, building software, and it already cost money to create. The licensing model earns a second return on that spend without changing what the company does. The full mechanics of the underlying agreement, from permitted uses to exclusivity to warranties, are in our guide to how AI data licensing deals work; this page is about the business decision above the contract.
The four ways it pays
| Model | How revenue works | Example |
|---|---|---|
| One-time archive license | A single fee for a finished corpus, sometimes paid in tranches | Spirit Airlines' operational history, a 10 million dollar bid from Google in bankruptcy court |
| Recurring license | Annual or multi-year fees for a living source, often with refreshes and API access | Google paying Reddit around 60 million dollars a year |
| Exclusive premium | A larger fee in exchange for denying the data to competitors for a term | Common in publisher and archive deals; the premium is priced against the doors it closes |
| Marketplace revenue share | The owner is paid as and when the data licenses through an aggregator | Thousands of creators and companies licensing through platforms such as Troveo |
The structures are not mutually exclusive. A publisher can run a recurring license with one lab and a non-exclusive archive license with another. A marketplace participant can license the same catalog to several buyers over time. Which structure fits depends on whether the data is finished or still being produced, how scarce it is, and how much optionality the owner wants to keep. The disclosed numbers behind each structure are collected on our AI training data statistics page, and the pricing factors are in our guide to what AI training data costs.
Examples across the spectrum
The model looks different at each scale, and the examples are the fastest way to see whether it applies to you.
At the top are media companies with famous catalogs. News Corp licensed its archive to OpenAI for a reported 250 million dollars plus over five years, and Shutterstock earned 104 million dollars from AI licensing in a single year. These are direct negotiations by owners with legal and data teams, and they set the public reference points for the whole market.
In the middle are platforms and communities whose value is that they keep producing. Reddit's deal with Google is the clearest example of the recurring model: a living source, refreshed continuously, paid for annually. Forums, review sites, and developer communities sit in the same category.
Then there are operating companies whose data was never meant to be a product. Spirit Airlines' archive became a 10 million dollar asset only in liquidation, and our breakdown of the Spirit Airlines data deal covers why it drew a bid at all: years of connected operational records across booking, operations, and finance. The lesson is not that every company's data is worth 10 million dollars. It is that operational history now has an observable market value, and that companies can engage with that market on their own terms rather than only in a wind-down. What that category includes is mapped in our guide to enterprise operational data.
Below that are startups shutting down. Closure platforms now broker code repositories, ticket histories, and workspace archives from defunct companies, with reported deals running from about 10,000 to 100,000 dollars each, and labs and research groups buy them to train agents on how real teams worked. How that corner of the market prices code is in our guide to selling source code to AI labs.
And at the widest point are individual creators and small companies licensing through marketplaces: video, audio, gameplay, and now business data, aggregated into one counterparty so that thousands of owners who could never negotiate with a lab directly can still be paid. Troveo alone has paid more than 20 million dollars to over 7,000 rights holders that way.
Why a business would choose it
The first reason is the one in the definition: the data already exists and already cost money to produce. Licensing it is revenue against a sunk cost, with no new product, no new headcount to serve customers, and no change to the core business. A non-exclusive license can earn that revenue more than once.
The second reason is control. Data that is not licensed does not stay unused. It gets scraped, it gets summarized by models trained on whatever was public, or, in the worst case, it gets auctioned by a trustee. Courts have concentrated liability on how data was acquired rather than on training itself, which is why buyers now pay for rights-cleared training data and why owners who license on their own terms set the scope, the exclusions, and the privacy standard instead of having those decided for them.
The third reason is that the market has matured enough to be worth the effort. Three years ago a mid-sized company had no realistic buyer for its operational records. Today there are published price signals, standard deal structures, aggregators that handle packaging and privacy review, and buyers actively looking for data that shows how real work gets done. The shift is visible in our guide to AI data licensing, which covers how licensing went from an afterthought to a standard procurement channel.
What it takes to run
A licensing model is a process, not a transaction, and the companies that do well at it treat it that way. The steps are consistent whatever the data: inventory what exists and where it lives; establish what the company owns and has authority to license; define scope, meaning what is in, what is excluded, and what safeguards apply; run privacy review and de-identification fit to the dataset; package the data so a buyer can evaluate it; then negotiate terms. Customer personal information, employee personal information, third-party confidential material, and anything under contractual restriction come out before anything ships. The owner's-side version of that process, step by step, is in our guide to licensing company data for AI.
The work is real, which is why most owners below publisher scale use a marketplace rather than building a data licensing function in-house. The aggregator handles buyer matching, standard terms, packaging, and privacy review once, at the platform level, and the owner decides what is in scope and gets paid as it licenses.
When it is the wrong model
Honesty about this is what separates a serious data strategy from a headline. Licensing is the wrong call when the data is not scarce: generic records that look like everyone else's have little training value. It is the wrong call when the rights are not clean and cannot be made clean, when customer contracts or regulation forbid it, or when the company cannot separate the operational knowledge in its data from the personal information of the people in it. It is the wrong call when an exclusive premium would close doors the company will want open in two years. And it is the wrong call when it distracts from the core business: for most companies this is a second revenue line, not a pivot. There is no universal price per record, and a company that cannot answer what makes its data different from a thousand others' should expect a modest outcome or none.
Does it apply to your company?
The fastest test is an inventory, and the shortcut to an inventory is Troveo's free data value assessment: nine questions about your systems, history, and industry, about five minutes, and a score for the AI-training value of what your company holds. It does not require a data team, and the result tells you whether a licensing conversation is worth having before you spend a day on one. The practical version of the question, what sells and what it pays, is in our guide to selling data to AI companies.
Where Troveo fits
Troveo is the marketplace version of the model, built for owners. It has spent years helping owners of proprietary real-world data license it for AI across video, audio, gaming, and robotics, with more than 20 million dollars paid to rights holders, and it now runs the same process for business data: identifying what is valuable, working with the owner on rights and scope, packaging for evaluation, and matching with AI buyers. Owners set the terms, exclusions come first, and payment follows licensing. Start with the data value assessment, or talk to us about whether a licensing model fits what your company holds.
