Simple AI Opens 2,000 Hours of Robot Training Data Collected Without Robots
The release makes a sizable human-capture corpus available for commercial use. Its early parity results are promising, but the robot-free system used more than 10 times as many demonstrations as teleoperation in key comparisons.
Story brief
3 key pointsSimple AI’s CC BY 4.0 release gives researchers a large, robot-independent source of bimanual demonstrations: 482,000-plus episodes across 110 scenes, with trajectories, video, language labels and validation metadata. In tests, policies trained on the data approached teleoperation results, including 85% success on precision insertion, but used over 10 times more demonstrations and benefited from shared gripper and...
- 01
HiFi-UMI-2K reports 2,000 hours, 482,000-plus episodes and approximately 98% reconstruction and simulation-replay validation pass rates.
- 02
Robot-free policies differed from teleoperation-trained results by -2.5, +3.1 and -0.6 percentage points across tested task groups.
- 03
The evaluation does not establish transfer to different grippers, sensor layouts, robot bodies or broader contact-heavy tasks.
Simple AI has released HiFi-UMI-2K, a 2,000-hour dataset designed to train robot arms from human demonstrations captured without a target robot or teleoperation rig. The experiment is not whether people can show a task, but whether those recordings are precise enough to replace part of the costly robot-operated data pipeline.
The dataset contains more than 482,000 episodes from over 110 scenes. It pairs synchronized multi-view video with two-handed trajectories, gripper states, language annotations, subtask boundaries and quality-control metadata. Simple AI released it under CC BY 4.0, allowing commercial use with attribution.
Capturing the action before a robot enters the picture
HiFi-UMI records people doing two-handed manipulation with portable, sensor-equipped grippers. The system then reconstructs those demonstrations as trajectories intended for real robot arms. That separates collection from the particular arm and control setup normally used in teleoperation.
What the capture system is built to preserve
- Head-mounted stereo cameras and inertial sensors record the operator’s view and movement.
- Each hand module uses two non-parallel fisheye cameras, while a shared hardware trigger aligns camera and inertial-sensor data.
- Simple AI reports roughly 3-millimeter local end-effector accuracy, synchronization below 40 microseconds and about 200 degrees of visual coverage around each hand.
After capture, the company reconstructs trajectories and tries to replay them in simulation, rejecting failures before they reach training. Simple AI reports about 98% pass rates for both trajectory reconstruction and simulation replay validation.
Near parity, but not equal data efficiency
Across three policy architectures and four bimanual tabletop task groups, policies trained only on HiFi-UMI data differed from teleoperation-post-trained policies by -2.5, +3.1 and -0.6 percentage points in aggregate success rate. On the strongest robot-free policy for a precision-insertion task, success reached 85%.
Those results show that high-fidelity human capture can work in the tested setup. They do not establish that it delivers the same performance from the same number of examples: the robot-free side had far more demonstrations. The teleoperation baseline was also collected in the evaluation scene, while the human demonstrations came from other locations.
A hardware boundary remains
The evaluation robot had different arm kinematics from the capture setup, but shared its gripper and wrist-camera configuration. The tests therefore do not yet answer how well the approach transfers to sharply different hands, sensor layouts or robot bodies, or to a wider range of contact-heavy tasks.
The public release is only a slice of the operation Simple AI describes. The company says it has processed a broader corpus of more than 20,000 hours and 4.32 million episodes across over 480 scenes. It also reports that pre-training on 4,000 hours from that corpus reduced offline action-prediction error by 41% across 10 unseen tasks and lifted real-robot success by 18.1 percentage points in its evaluated setup.
HiFi-UMI-2K gives researchers a commercial-use dataset to test that proposition themselves. Its central open question is economic as much as technical: whether capturing many more human demonstrations is cheap and fast enough to outweigh the lower per-trajectory efficiency suggested by the comparison.