世界模型有触觉了!50万小时视频,训出首个隐式触觉世界动作模型
Merged summary
TL;DR - BeingBeyond unveiled Being-H0.8, a tactile world-action model that learns contact-aware robot control from large-scale human video and sensor data. It matters because tactile prediction and feedback enable more adaptive, precise manipulation than vision-only control.
- Its training pipeline draws on over 500,000 hours of first-person video, with TactoHand generating contact and proximity supervision from footage lacking tactile labels.
- A universal tactile encoder standardizes heterogeneous contact, proximity, and pressure signals into fixed-format tokens; TopoHand aligns human hands, dexterous robot hands, and grippers.
- The model predicts interaction consequences in latent space rather than reconstructing future pixels.
- A slow-fast controller maintains the overall plan while repeatedly using fresh tactile feedback to regenerate short action segments during execution.
Sources (1)
世界模型有触觉了!50万小时视频,训出首个隐式触觉世界动作模型
TL;DR - BeingBeyond unveiled Being-H0.8, a tactile world-action model that learns contact-aware robot control from large-scale human video and sensor data. It matters because tactile prediction and feedback enable more adaptive, precise manipulation than vision-only control.
- Its training pipeline draws on over 500,000 hours of first-person video, with TactoHand generating contact and proximity supervision from footage lacking tactile labels.
- A universal tactile encoder standardizes heterogeneous contact, proximity, and pressure signals into fixed-format tokens; TopoHand aligns human hands, dexterous robot hands, and grippers.
- The model predicts interaction consequences in latent space rather than reconstructing future pixels.
- A slow-fast controller maintains the overall plan while repeatedly using fresh tactile feedback to regenerate short action segments during execution.