AI Training Data Glossary
Plain-language definitions for the licensing, rights, and provenance terms used across AI training data.
C
Chain of rights
The documented trail showing who owns a piece of content and that each link, from creator to licensor to buyer, had the authority to grant the rights it granted. In AI training data, a complete chain of rights per asset is what separates rights-cleared data from data that merely has a license attached.
Read definition →Content provenance (C2PA)
An industry standard for attaching tamper-evident origin records to media files, documenting where content came from and how it's been altered. In training data, provenance standards support the per-asset documentation that rights-cleared licensing requires.
Read definition →D
Data annotation
The work of labeling data so models can learn from it: tagging objects, transcribing speech, ranking outputs. Annotation companies process data you already have; they are a different category from licensing marketplaces, which supply the data itself.
Read definition →Data curation
The selection, cleaning, deduplication, and structuring of raw content into a training-ready dataset. Curation is where most of the practical value in a dataset gets added, and it's the difference between buying content and buying training data.
Read definition →Data exclusivity
A licensing term under which a rights holder agrees not to license the same content to other buyers. Exclusive data is the only data that differentiates a model, since content every lab can buy moves no benchmarks. On Troveo, 95 percent of licensors have signed exclusively.
Read definition →Data provenance
The documented origin and history of a piece of training data: who created it, who has held the rights to it, and how it was acquired. Courts in the AI copyright cases have treated acquisition as a separate legal question from training, which has made provenance a requirement rather than a preference for production models.
Read definition →E
F
Fair use
The legal doctrine allowing limited use of copyrighted material without permission, weighed on factors including the purpose of the use and its effect on the market for the original. In AI training, courts have split on when training qualifies, and the growth of the data licensing market itself weakens the defense over time, because a market that exists can be harmed.
Read definition →Fine-tuning data
Smaller, targeted datasets used to adapt a pretrained model to a specific task, domain, or behavior. Because fine-tuning data steers the model directly, quality and specificity matter more per example than in pretraining, which makes it a market for curated, licensed datasets rather than bulk collection.
Read definition →I
M
Model weights
The numerical parameters a model learns during training, the artifact that training data ultimately becomes. Weights matter in licensing because content absorbed into them cannot be simply deleted later, which is why training rights, unlike display rights, are effectively permanent grants.
Read definition →Multi-camera video
Footage of the same scene captured simultaneously from multiple angles. It teaches models spatial consistency, how a scene holds together across viewpoints, which makes it one of the most requested data types for world models and 3D-aware video generation.
Read definition →R
S
Sim-to-real gap
The performance drop that occurs when a model trained in simulation is deployed in the physical world, caused by everything simulators fail to reproduce. It is the main reason physical AI teams layer real-world footage on top of synthetic environments.
Read definition →Synthetic training data
Training data generated by models rather than collected from the world. It fills gaps and covers edge cases cheaply, but inherits the limitations of the models that generate it, so labs treat it as a layer on top of real-world data rather than a replacement.
Read definition →T
Teleoperation
Humans remotely controlling robots to perform tasks, recorded to produce robot-native training data. It yields the highest-fidelity demonstrations and costs the most per hour, which is why teams reserve it for the last mile and layer licensed real-world footage underneath.
Read definition →Training-ready data
Data delivered in the condition a training pipeline actually needs: cleaned, normalized, consistently formatted, with metadata and per-asset rights documentation. The difference between raw content and training-ready data is weeks of engineering and legal review, which is why buyers increasingly treat delivery condition as part of the price.
Read definition →