Industry analysis8 min read

Robot dataset acceptance: what to verify before scaling collection

An episode count is not a training-ready deliverable. Accept the task contract, robot configuration, multimodal timing, outcome labels, coverage and dataset revision before paying to scale collection.

By Matrix Dimension Robotics engineering team

Concept rendering of a wheeled humanoid robot at an industrial workstation illustrating embodied-AI data collection and acceptance

The short answer: do not buy robot demonstration data by episode count alone. First accept whether every episode can be interpreted, replayed, filtered, trained and tied to an independent evaluation. China's Ministry of Industry and Information Technology has placed Technical Requirements for Humanoid Robot Data Collection in its 2026 industry-standard development plan. That is a standards project, not a published or effective standard, but it is a useful signal to turn data collection into a controlled engineering deliverable now.

What the Chinese standards project does—and does not—mean

The official MIIT third batch of 2026 industry-standard projects lists project 2026-0576T-SJ, Technical Requirements for Humanoid Robot Data Collection. It is a new 12-month drafting project under the ministry's humanoid-robot and embodied-intelligence standardisation committee. The plan does not publish acceptance clauses. Procurement and review documents should therefore say “standards project established,” not “standard implemented,” “compliant” or “certified.”

The practical signal travels beyond humanoids and beyond China: robot data is becoming a product with a declared object, process, quality boundary and configuration history. Teams can prepare without guessing future clauses by making today's datasets traceable and replayable.

Six acceptance layers: volume is only the outer shell

LayerAcceptance questionMinimum evidenceCommon failure
Task and outcomeWhat starts, succeeds, fails or aborts the task?Task ID, instruction, initial state, terminal conditions and outcome labelReviewers assign different results to the same trajectory
Embodiment and configurationWhich robot and controlled configuration produced it?Robot, tool, sensors, mounting, calibration, controller and software revisionsCamera or end effector changes while old frames and normalisation remain
Modalities and semanticsWhat do image, state, action and force fields mean?Names, types, shapes, units, frames, sample rates and missing-value rulesMatching dimensions hide different units or action references
Timing and episode integrityDo the streams describe the same physical event?Clock source, timestamps, alignment, drops, latency and start/stop reasonsVideo and action are offset although playback looks complete
Quality and coverageDoes collection cover reachable deployment states?Scene, object, operator, disturbance and outcome distributions; review and exclusion rulesMany episodes repeat one camera pose, object and clean success path
Splits and revisionCan training results be reproduced and compared?Train/validation/test rules, revision log, exclusions and dataset–checkpoint bindingNeighbouring trajectories or the same scene leak into test

Readable is not yet acceptable

The official LeRobotDataset API is episode-aware and carries feature schemas, tasks, frame rate, episode metadata and statistics. Its dataset tools let a reviewer inspect camera streams, robot states and actions on a shared timeline. Those facilities are a strong format and review foundation. They do not, by themselves, prove that the task definition is correct, the timing error fits the application, or the collected distribution represents the intended workplace.

Quality does not mean deleting every failed attempt. LeRobot's human-in-the-loop workflow records interventions, corrections and recovery motions alongside autonomous segments to address states reached when policy errors compound. Research on data quality in imitation learning frames quality around distribution shift and shows that state diversity is not automatically beneficial in every setting. Failures and interventions need explicit classes and intended uses: neither silently mix them into demonstrations nor discard them indiscriminately.

A seven-step gate from pilot capture to released dataset

  1. Write the task contract. Freeze objects, tools, initial state, permitted actions, observable success, failure classes and abort conditions.
  2. Baseline one embodiment. Record tool, cameras, calibration, control rate and software for one robot. Never merge a changed configuration silently.
  3. Capture a small pilot. Prove integrity, alignment, replay, outcome review and exclusion before adding operators or stations.
  4. Use stratified review. Sample across operator, scene, task, outcome, configuration and disturbance—not only the newest or most attractive videos.
  5. Keep raw lineage. Every conversion, crop, annotation edit and exclusion should resolve back to its source episode and reason.
  6. Design the split before scaling. Hold out the station, object, collection batch or operator that represents the actual generalisation question. Do not divide one continuous session randomly across train and test.
  7. Release through independent rollouts. Bind dataset revision, training configuration, checkpoint, evaluation initial state and episode outcomes. Dataset QA passing is not policy deployment approval.

Coverage should answer a deployment question

The DROID project studies diverse real-world manipulation through distributed collection across scenes and tasks. It also released improved camera calibrations for part of the dataset in 2025. That update is a useful reminder: a dataset is not an immutable delivery folder. Calibration, annotation, exclusions and derived formats need explicit revision relationships. The LeRobot LIBERO documentation likewise asks result reports to pin the dataset revision and keep evaluation initial conditions controlled when comparing policies. Simulation specifics do not replace real-robot acceptance, but same revision, same conditions and reproducible comparison are equally valuable on hardware.

Boundary: this is a data-engineering preparation and delivery framework, not an interpretation of unreleased standard clauses. It makes no compliance, safety or performance claim for any dataset, policy or robot. Where recordings include people, location, speech, confidential process information or other identifiable data, the responsible parties must perform the separate contractual, legal, data-governance and security review applicable to their project.

Frequently asked questions

Does project 2026-0576T-SJ mean the Chinese standard is already effective?

No. It means Technical Requirements for Humanoid Robot Data Collection is in an industry-standard development plan. Its eventual text, designation, publication and effective status must be confirmed from later official notices.

How many robot demonstrations are enough?

There is no universal count independent of the task, embodiment, disturbances and target policy. Use a small pilot to prove schema, timing, labels and evaluation, then expand according to uncovered states and real-robot evaluation results.

Why not delete every failed trajectory?

Uncontrolled failed actions can contaminate demonstrations, but labelled deviations, interventions and recovery segments can define boundaries and teach recovery. Class and intended use matter more than a blanket keep-or-delete rule.

If LeRobot can load the dataset, is it training-ready?

No. Loading proves basic format compatibility. Acceptance still needs task semantics, units and frames, multimodal alignment, outcome labels, configuration lineage, coverage, split integrity and independent evaluation.

Sources

These primary sources support the material facts and engineering boundaries discussed above.

  1. 工业和信息化部 — 2026 年第三批行业标准制修订计划
  2. Hugging Face LeRobot — Dataset API and metadata
  3. Hugging Face LeRobot — Using Dataset Tools
  4. Hugging Face LeRobot — Human-in-the-Loop Data Collection
  5. Hugging Face LeRobot — LIBERO evaluation and dataset revisions
  6. DROID — A Large-Scale In-the-Wild Robot Manipulation Dataset
  7. Data Quality in Imitation Learning

Evaluating robot control, bimanual manipulation or a mobile platform?

Talk to Matrix Dimension →