Twelve episodes from our bimanual kitting cell: four expert demos and eight on-policy rollouts of a fine-tuned π0.5 policy, labeled frame by frame. Raw data is in the download at the bottom.
put the bolt in the bag — 19.0 s. The policy runs 4.5 seconds and stalls; the operator takes over, recovers it, and hands back twice before the policy finishes. Every frame carries this flag, plus a per-frame time-to-success target. Click the timeline to seek.
collect expert demos → fine-tune → run rollouts autonomously → capture interventions, label failures and recoveries → retrain → repeat
Teleoperated demos that seed the fine-tune.
The fine-tuned policy alone, start to finish.
Operator takeovers, flagged frame by frame. Orange is human control.
One failure with no rescue attempted, one where the rescue failed.
We trained a value function on this data’s time-to-success labels. Across 215 episodes, its predictions correlate 0.86 (median) with ground truth. Here it reads one rollout:
The prediction climbs on approach, crashes at the failed grasp, stays low through the takeover, and climbs to done once the policy finishes.
All twelve episodes: raw trajectories, all camera streams, metadata, README, and a manifest of every label.
Open the sample folder in Google Drive