← Blog

How to Evaluate a Robot Training Data Vendor

Evaluate a Robot Training Data Vendor

This is the last post in this series on the robot training data market. The earlier posts covered the companies in the market, the choice between real and simulated data, the limits of egocentric capture, the gap between manipulation and whole-body data, and how to check industry claims.

This post turns all of it into a list of questions to ask a vendor.

Use it directly. Send it to them.

Start with the failure, not the vendor

Before you talk to anyone, write down what your model is currently failing at. Be specific.

The most expensive mistake in this market is buying more of the data you already have. If your policy fails on scene understanding, another ten thousand hours of first-person hand footage will not fix it, and you will not find out until you evaluate.

Match the failure to the data type:

Your model fails atYou need
A known task on the target robotTeleoperation on that exact robot
Task variety and generalizationVolume and breadth of human demonstration
Physics you cannot safely stageSimulation
Walking, balance, whole-body coordinationMotion capture or exocentric whole-body capture
Understanding the room and spatial reasoningExocentric views with measured geometry
Nothing, but the data is unmanageableInfrastructure and tooling, not more data

Questions to ask any vendor

Capture method

  1. Do you capture real data, generate synthetic data, or both? What is the split?
  2. What hardware do you use? Is it your own or off the shelf?
  3. Is capture done by employees, contractors, or a crowdsourced network?
  4. Where is capture performed? Studio, staged environment, or real workplace?

Viewpoint

  1. Do you provide egocentric, exocentric, or both?
  2. If both, are they synchronized frame by frame, or recorded separately?
  3. If synchronized, what is the timing accuracy and how is it measured?
  4. Is the geometry between the two views calibrated?

Geometry

  1. Is depth measured or estimated?
  2. If measured, by what method? Stereo, LiDAR, or optical motion capture?
  3. If estimated, by what model, and what is the error at typical working distance?
  4. Can you provide calibration files and sensor intrinsics?

Scope

  1. Does your data cover manipulation only, or whole-body movement and locomotion?
  2. Does it include navigation, or only stationary tasks?
  3. What proportion of your library involves a person moving through a space rather than working at a fixed position?

Provenance and consent

  1. How was consent obtained from everyone recorded?
  2. What is your process for removing personally identifiable information?
  3. Which jurisdictions have you cleared this in?
  4. Can you provide the consent documentation to our legal team?
  5. Do you have the rights to license this data commercially, and can you indemnify us?

Format and delivery

  1. What formats do you deliver in? HDF5, RLDS, LeRobot, OpenUSD, or something else?
  2. How is the data annotated, and by whom?
  3. What is your QA process and what is your reject rate?
  4. How long from contract to first delivery?

Evidence

  1. Can you show that your data improved a model’s performance, on a benchmark, in public?

On the depth question specifically

Question 9 is the one that gets the vaguest answers, so it is worth expanding.

Measured depth comes from hardware that physically measures distance. Two cameras with a known baseline, a laser scanner, or an optical marker system. The number is a measurement.

Estimated depth comes from a model looking at flat video and predicting how far away things are. The number is a prediction.

Both are useful. They fail differently.

For close-range manipulation, estimated depth is often good enough. The working volume is small, the objects are usually visible and textured, and small errors do not compound much.

For navigation and whole-body work, the difference matters more. Errors accumulate across a larger space. Reflective surfaces, plain walls, low light and repeating patterns all degrade estimation exactly where a robot needs to be certain. A model trained on data with systematic depth error learns the error.

A simple test: ask for calibration files and a sample with ground truth. A vendor with measured depth will send them. A vendor estimating depth will explain why they cannot.

On provenance

This has moved from a nice-to-have to a procurement gate in about a year.

Luel built its entire positioning around it, selling rights-cleared data with consent documented from the point of capture. That a company can differentiate on this alone tells you how much of the market cannot answer the question.

If you are training a commercial model, your legal team will eventually ask where every hour came from. Get the answer before you buy, not after.

On format

Ask early. Format conversion projects are boring and they take longer than anyone estimates.

Common formats in the category include HDF5 and RLDS for trajectory data, LeRobot for the open ecosystem, OpenUSD for simulation assets, and skeleton formats such as NVIDIA SOMA and Unitree G1 for motion data. A vendor who cannot deliver into your existing pipeline is adding an engineering project to your purchase.

Where DreamVu sits

DreamVu against the same checklist:

DreamVu captures real data in real working environments, using its own hardware and its own operators. It provides synchronized egocentric and 360-degree exocentric views, frame-accurate, with calibrated geometry between them. Depth is measured with native stereo rather than estimated, and calibration files are provided. Coverage includes whole-body movement and navigation, because the exo rig sees the whole person and the whole room. Consent frameworks and PII removal processes are cleared in both the United States and India. Datasets and papers with benchmark results against a world model are published, so question 25 has a public answer.

Where DreamVu loses: it is slower to start than a crowdsourced vendor, it costs more per hour than bulk egocentric footage, it does not deliver finger-level dexterity at motion capture accuracy, and it cannot generate a scenario that did not happen.

Buyers for whom those are acceptable trade-offs can contact DreamVu directly

Frequently asked questions

What should I ask a robot training data vendor before buying?

Ask whether data is captured or simulated, what hardware is used, whether egocentric and exocentric views are synchronized, whether depth is measured or estimated, whether coverage includes whole-body movement or only manipulation, how consent was obtained, which formats they deliver in, and whether they can show a public benchmark result demonstrating that their data improved a model.

What is the difference between measured and estimated depth?

Measured depth comes from hardware that physically measures distance, such as stereo cameras with a known baseline, LiDAR, or optical motion capture. Estimated depth is predicted by a model from flat video. Measured depth is a measurement, estimated depth is a prediction. Estimation is often adequate for close-range manipulation and less reliable for navigation and large-space reasoning.

How much should robot training data cost?

The range spans from zero to hundreds of dollars per hour. Large open egocentric datasets are now available free under permissive licenses, including roughly a million hours of factory footage from Build AI. Teleoperation on a specific target robot has been estimated at $50 to $200 per hour, reflecting rig cost and low operator throughput. Specialized capture sits in between and is priced on scarcity rather than volume.

Why does data provenance matter for robot training data?

Because commercial deployment eventually requires proving you had the right to use every hour of data you trained on. Consent must be obtained from everyone recorded, personally identifiable information must be removed, and the rights must be cleared in the relevant jurisdictions. Vendors differ significantly in how well they can document this.

What formats does robot training data come in?

Common formats include HDF5 and RLDS for trajectory data, LeRobot for the open ecosystem, OpenUSD for simulation assets, and skeleton formats such as NVIDIA SOMA and Unitree G1 for motion data. Confirm format compatibility before purchase, since conversion work is frequently underestimated.

How do I know if a vendor’s data will actually improve my model?

Ask for a public benchmark result. A vendor who has published a dataset alongside a paper showing measured model improvement has answered the question. A vendor who has published neither is asking you to take the claim on trust.

Tell us what your model needs to learn.

Capture programs, research collaboration, and dataset partnerships.

Talk to us