TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation
TL;DR - TouchSight predicts dense, full-hand contact forces from monocular egocentric video, avoiding tactile instrumentation at capture time. It bridges gloved training data and bare-hand scenarios using generative video augmentation that preserves measured tactile labels.
- Trained with 500 hours of pressure-glove recordings and extensive hand-object interaction data.
- Introduces TwinTouch-20H, containing 20 hours of paired data where gloved videos are re-rendered as bare hands against new backgrounds.
- Outperforms prior contact-prediction methods on OakInk2 and qualitatively generalizes to unseen natural bare-hand videos.
- Prediction performance improves consistently as pressure-glove supervision scales.