Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
TL;DR - Facet-0 is a robotic foundation model for contact-rich, sub-millimeter assembly that jointly predicts actions and their expected force/torque consequences. It achieves 82% mean success across five computer-assembly tasks, versus 15% for the strongest baseline.
- Aligns vision-language semantics and robot kinematics with causal wrist-wrench histories.
- Uses flow matching to jointly generate action chunks and predicted future wrist-wrench profiles.
- Applies deployment-based RL with an Action-Wrench Critic, phase-aware rewards, and contact-selective credit assignment.
- Trained on the 1,000-hour, force-synchronized ManuFacet-1K dataset; reaches 0.5 mm placement accuracy at 50 ms command latency.