🛰️ Daily AI Frontier
‹ back to 2026-08-12

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Research Robot Learning

Ranking

Overall 66
Content 80
Popularity 33

Observed public metrics from 1 member.

Merged summary

TL;DR - Surgical WAM is a unified world-action model (built on Cosmos Policy) that pretrains on cheap, action-free endoscopic video and then fine-tunes on a fixed budget of action-labeled dVRK demonstrations, showing that video dynamics priors substantially improve closed-loop surgical manipulation. It matters because action-labeled surgical teleoperation data is the main bottleneck for scaling surgical robot learning.

  • Jointly predicts future endoscopic observations and executable action chunks, unlike prior surgical world models that use video only for simulation or policy evaluation rather than control.
  • Deployed as a receding-horizon closed-loop controller: executes a short prefix of each predicted action chunk, then replans from the resulting observation.
  • Across four simulated surgical manipulation tasks, action-free video pretraining raised average success rate from 63.5% to 77.8%, with a 20-point absolute gain on PegTransfer.
  • Gains were largest on contact-rich and bimanual tasks, supporting the claim that video supplies transferable visual dynamics priors under limited action supervision.

Sources (1)

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

arXiv cs.RO Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang 2026-08-11 arXiv:2608.11204
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-10 14:31:05.136976 UTC

TL;DR - Surgical WAM is a unified world-action model (built on Cosmos Policy) that pretrains on cheap, action-free endoscopic video and then fine-tunes on a fixed budget of action-labeled dVRK demonstrations, showing that video dynamics priors substantially improve closed-loop surgical manipulation. It matters because action-labeled surgical teleoperation data is the main bottleneck for scaling surgical robot learning.

  • Jointly predicts future endoscopic observations and executable action chunks, unlike prior surgical world models that use video only for simulation or policy evaluation rather than control.
  • Deployed as a receding-horizon closed-loop controller: executes a short prefix of each predicted action chunk, then replans from the resulting observation.
  • Across four simulated surgical manipulation tasks, action-free video pretraining raised average success rate from 63.5% to 77.8%, with a 20-point absolute gain on PegTransfer.
  • Gains were largest on contact-rich and bimanual tasks, supporting the claim that video supplies transferable visual dynamics priors under limited action supervision.
item →