Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - Surgical WAM is a unified world-action model (built on Cosmos Policy) that pretrains on cheap, action-free endoscopic video and then fine-tunes on a fixed budget of action-labeled dVRK demonstrations, showing that video dynamics priors substantially improve closed-loop surgical manipulation. It matters because action-labeled surgical teleoperation data is the main bottleneck for scaling surgical robot learning.
- Jointly predicts future endoscopic observations and executable action chunks, unlike prior surgical world models that use video only for simulation or policy evaluation rather than control.
- Deployed as a receding-horizon closed-loop controller: executes a short prefix of each predicted action chunk, then replans from the resulting observation.
- Across four simulated surgical manipulation tasks, action-free video pretraining raised average success rate from 63.5% to 77.8%, with a 20-point absolute gain on PegTransfer.
- Gains were largest on contact-rich and bimanual tasks, supporting the claim that video supplies transferable visual dynamics priors under limited action supervision.
Sources (1)
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
TL;DR - Surgical WAM is a unified world-action model (built on Cosmos Policy) that pretrains on cheap, action-free endoscopic video and then fine-tunes on a fixed budget of action-labeled dVRK demonstrations, showing that video dynamics priors substantially improve closed-loop surgical manipulation. It matters because action-labeled surgical teleoperation data is the main bottleneck for scaling surgical robot learning.
- Jointly predicts future endoscopic observations and executable action chunks, unlike prior surgical world models that use video only for simulation or policy evaluation rather than control.
- Deployed as a receding-horizon closed-loop controller: executes a short prefix of each predicted action chunk, then replans from the resulting observation.
- Across four simulated surgical manipulation tasks, action-free video pretraining raised average success rate from 63.5% to 77.8%, with a 20-point absolute gain on PegTransfer.
- Gains were largest on contact-rich and bimanual tasks, supporting the claim that video supplies transferable visual dynamics priors under limited action supervision.