🛰️ Daily AI Frontier
‹ back to 2026-08-07

无本体数据直达真机,深朴智能拔掉机器人后训练的「真机数据锚点」

Industry & News Embodied AI Data

Ranking

Overall 70
Content 70
Popularity 70

Observed public metrics from 1 member.

Representative image for 无本体数据直达真机,深朴智能拔掉机器人后训练的「真机数据锚点」

Merged summary

TL;DR — 深朴智能 (Simple AI) released HiFi-UMI, a high-fidelity handheld ("bodiless"/UMI-style) data-collection and production engine for robot manipulation, claiming that post-training on HiFi-UMI demonstrations alone can be deployed to real robots without any in-domain teleoperation "anchor" data.

  • Hardware/pipeline redesign targets four fidelity gaps in classic UMI: head-mounted stereo+IMU offline SLAM (~3 mm end-effector trajectory error, no external mocap), native bimanual relative pose via head-camera-tracked hand markers, unified GPIO hardware sync (<40 µs cross-sensor skew), and 6 cameras (2 head + 2 ultra-wide fisheye per wrist, ~200° FoV) plus a full-palm glove gripper.
  • Real-robot results across 3 models, 2 architectures: on 4 dual-arm tasks (wiping, shirt folding, remote-control insertion, fruit/veg sorting), 40 trials per task-policy pair. StarVLA-QwenPI 51.3% vs 53.8% teleop (−2.5%), OpenPI-π0.5 77.5% vs 74.4% (+3.1%), LingBot-VA (WAM) 56.9% vs 57.5% (−0.6%) — despite HiFi-UMI data being collected out-of-domain while teleop baselines were in-domain.
  • Pretraining scaling: 4,000 hours of multi-task data cut mean action-prediction error 41% on 10 unseen tasks and raised the 4-task real-robot success rate a further 18.1%; error fell 61% overall with a power-law fit (α=0.268, R²=0.993). Pretrained models needed only 800 insertion demos to beat a non-pretrained model trained on 3,200.
  • Ecosystem/positioning: pipeline has processed >20,000 hours / 4.32M segments across 480+ scenes; a curated 2,000-hour subset (HiFi-UMI-2K) is open-sourced on Hugging Face with synced six-view video, calibrated bimanual trajectories, gripper state, language descriptions and sub-task boundaries. Transfer correlated with coverage of interaction type (rigid pick-place >1/3 of frames, improved fastest) rather than object novelty; deformable folding (<1% of frames) improved least.

Sources (1)

无本体数据直达真机,深朴智能拔掉机器人后训练的「真机数据锚点」

WeChat: 机器之心 2026-08-04 arXiv:2607.25895
Public signals Hugging Face upvotes 158
Providers: Hugging Face · Upvotes 158 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-03 14:30:24.819361 UTC

TL;DR — 深朴智能 (Simple AI) released HiFi-UMI, a high-fidelity handheld ("bodiless"/UMI-style) data-collection and production engine for robot manipulation, claiming that post-training on HiFi-UMI demonstrations alone can be deployed to real robots without any in-domain teleoperation "anchor" data.

  • Hardware/pipeline redesign targets four fidelity gaps in classic UMI: head-mounted stereo+IMU offline SLAM (~3 mm end-effector trajectory error, no external mocap), native bimanual relative pose via head-camera-tracked hand markers, unified GPIO hardware sync (<40 µs cross-sensor skew), and 6 cameras (2 head + 2 ultra-wide fisheye per wrist, ~200° FoV) plus a full-palm glove gripper.
  • Real-robot results across 3 models, 2 architectures: on 4 dual-arm tasks (wiping, shirt folding, remote-control insertion, fruit/veg sorting), 40 trials per task-policy pair. StarVLA-QwenPI 51.3% vs 53.8% teleop (−2.5%), OpenPI-π0.5 77.5% vs 74.4% (+3.1%), LingBot-VA (WAM) 56.9% vs 57.5% (−0.6%) — despite HiFi-UMI data being collected out-of-domain while teleop baselines were in-domain.
  • Pretraining scaling: 4,000 hours of multi-task data cut mean action-prediction error 41% on 10 unseen tasks and raised the 4-task real-robot success rate a further 18.1%; error fell 61% overall with a power-law fit (α=0.268, R²=0.993). Pretrained models needed only 800 insertion demos to beat a non-pretrained model trained on 3,200.
  • Ecosystem/positioning: pipeline has processed >20,000 hours / 4.32M segments across 480+ scenes; a curated 2,000-hour subset (HiFi-UMI-2K) is open-sourced on Hugging Face with synced six-view video, calibrated bimanual trajectories, gripper state, language descriptions and sub-task boundaries. Transfer correlated with coverage of interaction type (rigid pick-place >1/3 of frames, improved fastest) rather than object novelty; deformable folding (<1% of frames) improved least.
item →