Robotics has a data problem the rest of AI already solved once. Language models had the web to pretrain on. Robot foundation models have nothing equivalent: there is no internet of robot experience, and physical interaction data has to be captured, generated, or licensed deliberately. That scarcity has produced a vendor ecosystem in a hurry, and buyer questions about it show up constantly: which vendor should a robotics startup use, who combines simulation with real-world data, who actually collects demonstrations. This guide maps who does what.
If you want the primer on what this data is before the vendor list, our explainer on robotics training data covers the data types themselves.
What robotics teams actually buy
Physical AI programs stack several kinds of data. Teleoperation and demonstration data, where humans operate or guide robots through tasks, is the highest-signal and most expensive layer. Egocentric human video, footage of people doing tasks from a first-person or body-worn perspective, teaches models how the physical world works before a robot ever moves. Simulation provides scale and edge cases that would be dangerous or slow to capture physically. RL environments turn all of it into training loops for agentic behavior. And annotation infrastructure sits underneath, curating and labeling what gets collected.
Almost no vendor covers the whole stack, which is why buyers end up comparing companies that do completely different things.
The companies
At the full-stack end, Scale AI runs physical AI data programs spanning human demonstration capture and dedicated collection operations, the closest thing to a one-stop shop, with the Meta ownership caveat that applies across its business; our guide to Scale AI alternatives covers that tradeoff. Centific and Toloka both run managed global workforces that have expanded toward physical AI collection, and they show up together in buyer comparisons for robotics and sensor data work; our guide to Centific alternatives maps that services field.
The specialist layer is younger and moving fast. Teleoperation and demonstration startups build capture rigs and sell bimanual manipulation datasets. Egocentric capture companies deploy head-mounted and body-worn rigs in homes and factories. On the simulation side, NVIDIA's Isaac and Omniverse ecosystem is the gravitational center, with startups building physics-accurate assets and environments on top of it. Most of these specialists are small, recently funded, and worth evaluating case by case, since capability claims in this segment move faster than delivered datasets.
For the agentic training layer, a separate market builds RL environments as a service, covered in full in our guide to RL environment companies. And underneath it all, annotation platforms like Encord handle curation, versioning, and labeling of robot data.
The layer most teams discover last is licensed real-world video: footage of humans performing tasks, manipulating objects, and navigating spaces, captured at a scale no rig deployment can match. That is where marketplaces enter the picture, including Troveo.
The landscape
| Provider | Category | Best for |
|---|---|---|
| Scale AI | Full-stack physical AI data | Large programs, demonstration capture, Meta caveat applies |
| Centific | Managed data services | Multilingual workforce, sensor and collection programs |
| Toloka | Managed global crowds | Distributed physical AI data collection |
| NVIDIA (Isaac, Omniverse) | Simulation ecosystem | Physics-accurate synthetic data at scale |
| Teleoperation specialists | Demonstration capture | Bimanual manipulation and rig-based datasets |
| Egocentric capture startups | First-person collection | Purpose-built home and factory footage |
| Encord | Annotation infrastructure | Curation and labeling of robot datasets |
| RL environment providers | Agentic training loops | Turning data into trainable environments |
| Mercor | Expert marketplace | Human experts for evals and demonstrations |
| Troveo | Licensed data marketplace | Rights-cleared real-world task, manipulation, and navigation video |
The underrated layer: licensed real-world video
Robot learning research keeps converging on the same finding: models pretrain on human video before they fine-tune on robot data. Humans manipulating tools, interacting with objects, moving through homes, warehouses, and streets, that footage encodes the physics, affordances, and task structure a robot needs, and it already exists in the world at a volume no collection program can replicate. The constraint is rights: scraped footage is a legal liability, and commissioned capture is slow and expensive per hour.
That is the gap licensed marketplaces fill. Troveo's catalog includes task demonstrations, tool manipulation, human-object interaction, open-world navigation, and first-person footage, sourced from more than 7,000 rights holders and cleared for AI training with documentation per asset. The same footage that feeds gameplay and world model training serves physical AI teams building world models for robots.
How to choose
Match the vendor to the layer you are missing. If you have no demonstration data, that is a teleoperation or full-stack conversation. If you have demonstrations but no scale, that is simulation plus licensed video. If you have data but no training loop, that is the RL environment market. Then apply the standard checks: ask specialists for delivered-dataset references rather than capability claims, ask simulation vendors how sim-to-real gaps are measured, and ask anyone selling real-world footage where the rights come from, per asset, because provenance problems in robotics data are the same lawsuits waiting to happen as everywhere else in AI.
Where Troveo fits
Troveo is the licensed data layer of this stack: real-world video of humans doing the things robots are learning to do, at a scale that rig-based collection cannot reach, with the rights documentation physical AI teams need. Teams combine it with their own robot data, simulation, and RL environments. Browse task demonstration, manipulation, and navigation datasets in Lens or talk to us about what your robots are learning.
