AI Training Data Glossary

Plain-language definitions for the licensing, rights, and provenance terms used across AI training data.

C

D

Dark data

Dark data is information an organization has collected or generated but is not actively using for its primary business purpose: old email archives, logs, closed tickets, and records kept for compliance. It used to be a storage cost. AI changed that, because some proprietary dark data contains real operational knowledge that exists nowhere on the public web. Not all of it is valuable; worth depends on uniqueness, connected context, visible outcomes, and whether the rights are clean enough to license.

Read definition →

Data annotation

The work of labeling data so models can learn from it: tagging objects, transcribing speech, ranking outputs. Annotation companies process data you already have; they are a different category from licensing marketplaces, which supply the data itself.

Read definition →

Data curation

The selection, cleaning, deduplication, and structuring of raw content into a training-ready dataset. Curation is where most of the practical value in a dataset gets added, and it's the difference between buying content and buying training data.

Read definition →

Data exclusivity

A licensing term under which a rights holder agrees not to license the same content to other buyers. Exclusive data is the only data that differentiates a model, since content every lab can buy moves no benchmarks. On Troveo, 95 percent of licensors have signed exclusively.

Read definition →

Data provenance

The documented origin and history of a piece of training data: who created it, who has held the rights to it, and how it was acquired. Courts in the AI copyright cases have treated acquisition as a separate legal question from training, which has made provenance a requirement rather than a preference for production models.

Read definition →

E

F

I

L

M

R

S

T

V

W

Browse all resources →