无本体数据直达真机,深朴智能拔掉机器人后训练的「真机数据锚点」
Ranking
Overall
70
Content
70
Popularity
70
Observed public metrics from 1 member.
Merged summary
TL;DR — 深朴智能 (Simple AI) released HiFi-UMI, a high-fidelity handheld ("bodiless"/UMI-style) data-collection and production engine for robot manipulation, claiming that post-training on HiFi-UMI demonstrations alone can be deployed to real robots without any in-domain teleoperation "anchor" data.
- Hardware/pipeline redesign targets four fidelity gaps in classic UMI: head-mounted stereo+IMU offline SLAM (~3 mm end-effector trajectory error, no external mocap), native bimanual relative pose via head-camera-tracked hand markers, unified GPIO hardware sync (<40 µs cross-sensor skew), and 6 cameras (2 head + 2 ultra-wide fisheye per wrist, ~200° FoV) plus a full-palm glove gripper.
- Real-robot results across 3 models, 2 architectures: on 4 dual-arm tasks (wiping, shirt folding, remote-control insertion, fruit/veg sorting), 40 trials per task-policy pair. StarVLA-QwenPI 51.3% vs 53.8% teleop (−2.5%), OpenPI-π0.5 77.5% vs 74.4% (+3.1%), LingBot-VA (WAM) 56.9% vs 57.5% (−0.6%) — despite HiFi-UMI data being collected out-of-domain while teleop baselines were in-domain.
- Pretraining scaling: 4,000 hours of multi-task data cut mean action-prediction error 41% on 10 unseen tasks and raised the 4-task real-robot success rate a further 18.1%; error fell 61% overall with a power-law fit (α=0.268, R²=0.993). Pretrained models needed only 800 insertion demos to beat a non-pretrained model trained on 3,200.
- Ecosystem/positioning: pipeline has processed >20,000 hours / 4.32M segments across 480+ scenes; a curated 2,000-hour subset (HiFi-UMI-2K) is open-sourced on Hugging Face with synced six-view video, calibrated bimanual trajectories, gripper state, language descriptions and sub-task boundaries. Transfer correlated with coverage of interaction type (rigid pick-place >1/3 of frames, improved fastest) rather than object novelty; deformable folding (<1% of frames) improved least.
Sources (1)
无本体数据直达真机,深朴智能拔掉机器人后训练的「真机数据锚点」
Public signals
Hugging Face upvotes 158
TL;DR — 深朴智能 (Simple AI) released HiFi-UMI, a high-fidelity handheld ("bodiless"/UMI-style) data-collection and production engine for robot manipulation, claiming that post-training on HiFi-UMI demonstrations alone can be deployed to real robots without any in-domain teleoperation "anchor" data.
- Hardware/pipeline redesign targets four fidelity gaps in classic UMI: head-mounted stereo+IMU offline SLAM (~3 mm end-effector trajectory error, no external mocap), native bimanual relative pose via head-camera-tracked hand markers, unified GPIO hardware sync (<40 µs cross-sensor skew), and 6 cameras (2 head + 2 ultra-wide fisheye per wrist, ~200° FoV) plus a full-palm glove gripper.
- Real-robot results across 3 models, 2 architectures: on 4 dual-arm tasks (wiping, shirt folding, remote-control insertion, fruit/veg sorting), 40 trials per task-policy pair. StarVLA-QwenPI 51.3% vs 53.8% teleop (−2.5%), OpenPI-π0.5 77.5% vs 74.4% (+3.1%), LingBot-VA (WAM) 56.9% vs 57.5% (−0.6%) — despite HiFi-UMI data being collected out-of-domain while teleop baselines were in-domain.
- Pretraining scaling: 4,000 hours of multi-task data cut mean action-prediction error 41% on 10 unseen tasks and raised the 4-task real-robot success rate a further 18.1%; error fell 61% overall with a power-law fit (α=0.268, R²=0.993). Pretrained models needed only 800 insertion demos to beat a non-pretrained model trained on 3,200.
- Ecosystem/positioning: pipeline has processed >20,000 hours / 4.32M segments across 480+ scenes; a curated 2,000-hour subset (HiFi-UMI-2K) is open-sourced on Hugging Face with synced six-view video, calibrated bimanual trajectories, gripper state, language descriptions and sub-task boundaries. Transfer correlated with coverage of interaction type (rigid pick-place >1/3 of frames, improved fastest) rather than object novelty; deformable folding (<1% of frames) improved least.