Physical AI · evidence and evaluation

Where the demonstrations came from, and what you are allowed to do with them

A robot dataset carries provenance and permission questions that a model card rarely answers: who was recorded, under which licence, and whether people appear in the frames.

How it works

Demonstration data is collected in real spaces, often with cameras that capture more than the task. Three questions decide whether it can be used: the licence of the dataset and of each constituent source; whether people, faces or private spaces appear and under what consent; and whether the recorded operators agreed to their behaviour being redistributed. Pooled datasets make this harder, because a permissive aggregate licence does not override the terms of a constituent subset. Maha keeps the detailed lineage and rights machinery in the robotics section and links to it rather than duplicating it here.

A concrete case

A policy is fine-tuned on an internal dataset recorded in a working lab. Before release, the licence question is not only about the images but about whether staff who appear incidentally in the wide-angle camera agreed to that use.

What this establishes

Nothing empirical. This is a rights and provenance checklist Maha proposes for teams building on demonstration data.

What it does not

Not legal advice. Licence interpretation and consent requirements depend on jurisdiction and on the specific agreements involved.

Questions worth asking

  • Ask for the licence of each constituent dataset, not only the aggregate.
  • Ask whether people appear in any frame and what consent covers that appearance.
  • Ask whether the deployment use is within the licence the data was collected under.

Sources

This page proposes a method rather than reporting a finding about the world, so it cites no external source. Where it describes something published, that page carries the citation.

Continue

Elsewhere on this site

Maha Strategies publishes explanation and evaluation method. We build no robots, run no physical experiments, and report no benchmark results of our own. Hardware, safety and evidence-intake questions live in the robotics section.