Guides12 min read

How Much Is Your Company's Data Worth to AI Labs? (What Buyers Pay in 2026)

Troveo Team

Troveo

Your company's data is worth somewhere between nothing and a life-changing number, and the honest answer to "how much is my data worth" is that nobody can tell you without seeing what you hold. What can be said is more useful than it sounds. The market now has public price points, from a few thousand dollars for a single code repository to tens of millions for a company archive. It has a floor: Troveo's business data page puts typical full-company deals at six figures. And it has a consistent set of factors that decide where a given company lands inside that range. This guide covers all three, explains how AI labs actually decide what to pay, and ends with the fastest way to get a real number for your own data instead of a guess.

Article banner reading What Is It Worth, on how much company data is worth to AI labs

The public price points

Most AI data deals are confidential, so the visible numbers are the edge of the market rather than its average. They are still the best reference points a company has, and they cluster by what was sold.

What was licensed or soldReported valueWhat it tells you
A single code repository (closure marketplace baseline)About 5,000 dollars per repoThe entry price for a small, clean, self-contained asset
Code and workspace archives from startups shutting downRoughly 10,000 to 100,000 dollars per dealWhat a few years of real engineering history is worth as a one-time sale
A full company's operational data (Troveo's business data program)Typical deals start at six figuresThe floor for a connected, multi-system, multi-year record, licensed rather than sold
Spirit Airlines' operational archive (Google's bid in bankruptcy court)10 million dollarsWhat a large company's complete history drew as a single flat purchase
Reddit's data access (Google)About 60 million dollars a yearA living source, priced as a recurring license
News Corp's catalog (OpenAI)More than 250 million dollars over five yearsThe top of the disclosed market: a famous, unique corpus with continuous updates
Public price points for data licensed or sold to AI developers, and what each one tells a company about its own data.

The closure-market figures come from reporting on platforms that broker assets from companies winding down, covered in detail in our guide to selling source code to AI labs. The Spirit figure is a court-record bid, and our breakdown of the Spirit Airlines data deal explains why an airline's internal records drew a bid at all. The publisher and platform deals, with sources, are kept current on our AI training data statistics page.

What the numbers tell you, and what they do not

The range is wide because the things being priced are different products, not different prices for the same thing. A repository is a finished, bounded asset. A company archive is years of connected activity across many systems. A publisher's catalog is a brand-name corpus that keeps growing. Reading the table as "data is worth 5,000 dollars" or "data is worth 60 million" is the mistake; the useful reading is that each row shows what a particular shape of data earned, and your company's data has a shape too.

The Spirit number is the one most often misread. It reflected an airline with tens of thousands of employees and decades of operations, sold in a liquidation where the buyer took the whole archive in one purchase. A 60-person software company is not going to see that figure. What Spirit established is narrower and more useful: operational history has an observable market value, and a company can engage with that market on its own terms while it is still running, rather than only when a trustee does it for them. The mechanics of that market, and why the structure of a deal moves the number as much as the data does, are in our guide to how AI data licensing deals work.

How AI labs actually value data

Buyers do not price data by the gigabyte, and they rarely start from a rate card. An AI lab evaluating a dataset is asking a narrower question: what does this let us train or test that we cannot already do with what we have? The answer sets the ceiling. Data that fills a specific gap, such as real multi-step business workflows with outcomes, real support conversations with resolutions, or real engineering history with the reviews and the bugs, is worth what closing that gap is worth to the buyer. Data that duplicates what is already on the public web is worth close to nothing, however much of it there is.

Three things follow from that. First, a buyer will usually want to evaluate a sample before committing, so packaging and documentation are part of the value, not overhead. Second, the same dataset can be worth very different amounts to different buyers, which is why marketplaces that reach many buyers tend to find higher prices than a single direct negotiation. Third, the pricing structure follows the data: a finished archive is bought flat, a living source is licensed on a recurring basis, and marketplace data is usually priced per unit or as a revenue share. The buyer's side of this, including what the disclosed deals reveal about pricing logic, is covered in our guide to AI training data cost.

The six factors that move a valuation

Inside the range, six factors do most of the work. They apply whether the company has 50 employees or 5,000.

Uniqueness. Data that exists nowhere else is the whole premise. A specialized product, an unusual industry, a process most companies do not run, or a customer base with distinctive behavior all push value up. A company whose records look like every other company's records pushes it down.

History. Five years beats five months. Buyers pay for patterns, and patterns need time: seasonal cycles, product changes, how a team's practices evolved, how problems recurred and were eventually solved. A long history is also harder to substitute, because nobody can generate it quickly.

Connected context. A support ticket on its own is a data point. The same ticket linked to the Slack thread where it was diagnosed, the Jira issue that tracked it, the pull request that fixed it, and the follow-up to the customer is a complete example of real work. Records that connect across systems are worth more than the sum of the exports, and Troveo's business data page says it plainly: the more systems licensed, and the more users and years behind them, the more valuable the data becomes. Our guide to enterprise operational data explains why buyers think in connected records rather than single sources.

Visible outcomes. A record of work with the result attached, resolved or escalated, won or lost, merged or rejected, is training material. The same record without the result is much less useful, because a model cannot learn what good looks like if it never sees what happened next. The glossary entry on workflow trajectories covers why the sequence with its outcome is the unit buyers care about.

Rights cleanliness. Buyers pay for licensed data specifically because it comes with documented, lawful provenance. Data the company clearly owns, with personal information removed and third-party restrictions respected, is worth more than data with an unclear chain of custody, and data the company cannot prove it has the right to license is worth nothing at all to a serious buyer. What that documentation looks like is in our guide to AI data provenance.

Buyer demand. The last factor is outside the company's control. Demand moves by category and by year: agent training pushed up the value of workflow and code data, and voice products did the same for real conversation. Knowing which categories are being bought right now is part of any valuation, and our guide to selling data to AI companies tracks what buyers are paying for.

Why volume is the wrong number to lead with

Founders tend to start with size: terabytes of logs, millions of rows, years of email. Volume matters far less than assumed. A thousand resolved support tickets with full context and outcomes are worth more than a million log lines, and a well-documented repository with five years of commits, reviews, and bug fixes is worth more than ten abandoned ones. Buyers who are training models to do real work want depth and connection, not bulk, and they can already get bulk from the public web.

This cuts the other way too. A company that assumes it is too small to matter is usually wrong. Value comes from uniqueness, history, and context rather than headcount, and a 40-person firm with a decade of specialized operational records can hold more training value than a 4,000-person company with generic data. What a typical startup stack holds, system by system, is in our guide to how tech startups and SaaS companies license their data to AI.

Licensing changes the math

Everything above describes what a dataset is worth in one transaction. Most companies do not have to settle for one transaction, and that changes the total more than any single factor.

Selling data is a one-time payment for something that could have earned more than once. Licensing keeps ownership with the company and grants a right to use a de-identified copy, which means the same dataset can be licensed to more than one buyer and turned into more than one product: a historical export, an enriched dataset with added structure, a continuing feed, and evaluation data built on top of it. Troveo's business data page describes the model as one dataset, many products, with the owner sharing in every sale, and gives four times the average sales per dataset as an illustrative example of how the total compounds. Exclusivity adds a further lever: an exclusive license commands a premium because the buyer gets something competitors cannot also have, and a non-exclusive one can be sold repeatedly. Which is worth more in a given case is worked through in our guide to exclusive vs non-exclusive data licenses, and the economics of running this as a revenue line rather than a single event are in our guide to data licensing as a business model.

The practical consequence is that "what is my data worth" has two answers: what a first deal pays, and what the data can earn over time. The first is the floor. The second is where a licensing program is judged.

What lowers the number

It is worth being direct about the other direction, because it saves companies from valuing the wrong thing. Data with no outcomes attached prices down. Generic data that looks like everyone else's prices down. Data the company cannot cleanly show it owns, or that sits under a customer contract restricting its use, prices at zero until that is resolved. And the data a company is often proudest of, the product usage data and customer records it collects, is frequently the least licensable, because customer contracts and privacy law restrict it regardless of de-identification. A company's most valuable licensable data is usually the record of its own work: engineering, support, sales process, operations, and decision records. That reframing alone changes most valuations, and our owner's guide to licensing company data for AI walks through the inventory and rights review that establish what is actually in scope.

How to get a real number

A serious valuation starts with an inventory, not a quote. List the systems, the years of history in each, the number of users, and the categories of records with outcomes. Establish what the company owns and what customer contracts or privacy obligations exclude. Then get the data in front of buyers, which for almost every company means going through a marketplace rather than cold-emailing labs.

The quickest first step is Troveo's free data value assessment: about five minutes, ten questions about the company's systems, history, and industry, and an estimate of what the data could license for. It is directional by design; a scoping call turns it into a real number, and Troveo's business data page is explicit that there are no fees, no setup costs, and no deductions from payouts, so finding out costs nothing but the time.

Where Troveo fits

Troveo helps companies understand what proprietary data they hold, protect what matters, and selectively license what is valuable. It has paid more than 20 million dollars to rights holders across video, audio, gaming, robotics, and business data, works with more than 40 active buyers, and runs the same process for operating companies: inventory, rights review, scoping, de-identification to a documented standard, packaging, and buyer matching, with the owner deciding what is in and what is out at every step and sharing in every sale rather than just the first. Start with the data value assessment to see what your company holds, or talk to us about what it might be worth.

Frequently asked questions

How much is my company's data worth to AI companies?
There is no rate card. Public reference points run from about 5,000 dollars for a single code repository to tens of millions for a large company's archive, and Troveo's business data page puts typical full-company deals at six figures. Where a specific company lands depends on uniqueness, history, connected context, outcomes, rights, and current buyer demand.
How do AI labs value training data?
By what it lets them train or test that they could not already do. Data that fills a specific gap, such as real workflows with outcomes or real engineering history, is priced against the value of closing that gap. Data that duplicates the public web is worth close to nothing regardless of volume. Buyers usually evaluate a sample before committing.
How much do AI companies pay for data?
Disclosed deals include roughly 10,000 to 100,000 dollars for startup code and archives, a 10 million dollar bid for Spirit Airlines' operational history, about 60 million dollars a year for Reddit's data, and more than 250 million dollars over five years for News Corp's catalog. Most deals are confidential, so these are the visible edge of the market.
Is a small company's data worth anything?
Often, yes. Value comes from uniqueness, years of history, and connected records with outcomes, not headcount. A small firm with a decade of specialized operational records can hold more training value than a large company with generic data. An inventory is how to find out.
How much is a codebase worth to an AI lab?
Reported closure-market deals run from about 5,000 dollars for a single repository to roughly 10,000 to 100,000 dollars for a startup's code and workspace archive. Value rises with history, review and bug-fix records, documentation, and clean rights, and falls with restrictive third-party licenses or missing context.
What makes company data worth less?
No outcomes attached, generic records that look like every other company's, unclear ownership, customer contracts that restrict use, and personal information that has not been removed. Product usage and customer data are often the least licensable parts of a company, while the record of its own work is usually the most valuable.
Do I get paid once or every time my data is used?
It depends on whether you sell or license. A sale is a single payment and the buyer takes ownership. A license keeps ownership with you and can be granted to more than one buyer or turned into more than one product, with the owner sharing in each sale. Troveo's business data program is built on the licensing model.
How do I find out what my data is actually worth?
Inventory the systems, years, users, and record types with outcomes; establish what you own and what is excluded; then get the data in front of buyers through a marketplace. Troveo's free data value assessment takes about five minutes and gives a directional estimate, and a scoping call turns it into a real number.

Related articles

Back to Resources