Robotics evidence guide · automated editorial preparation, not expert review
Keep training exposure out of robot evaluation claims
Record how evaluation episodes differ from training and tuning material. A different filename is not evidence of an independent test.
Evidence and interpretation
LeRobot’s metadata-based episode organization provides identifiers for an audit, but it does not certify a clean split. Maha proposes tracking collection sessions, source episodes and derived clips together. A training crop and an evaluation crop can come from the same demonstration. A held-out object can also appear in tuning footage under another label. Document what was checked and what cannot be known about upstream model training. Do not convert unavailable training provenance into a claim of no overlap.
Proposed evidence workflow
- Freeze train, tuning and evaluation episode identifiers together with their source lineage.
- Check shared sessions and derived media, not only exact file duplicates.
- Report known overlap, inspected non-overlap and unknown upstream exposure as different states.
Worked illustration
Ten clips from one synthetic session are divided eight for training and two for testing. The file split is valid bookkeeping but does not establish performance on a new session. A session-held-out claim is refused until a genuinely separate collection is used.
Limits
No leakage detector or model-training audit has been executed for external datasets. Unknown exposure remains unknown.
Sources and review
Locator: What’s new in v3; Format design
The format joins time-series signals and video through episode metadata; episode boundaries are not necessarily file boundaries.
Boundary: A storage format does not establish consent, dataset coverage, annotation accuracy or model generalization.
Inspection and reuse basis
Selected sections inspected 2026-09-14. Original explanation and link; no dataset or video redistribution. This is not independent expert review of this article.