Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Published
Source
arXiv
Paper number
882
Field
Robotics
arXiv ID
2608.11204

Key points

  • The authors propose a two-stage learning method that first learns visual dynamics in surgical scenes from unlabeled endoscopic video and then fine-tunes on a small amount of labeled action data.
  • Video pretraining raises the average success rate on four SurRoL benchmark tasks from 63.5 percent to 77.8 percent.
  • On PegTransfer, it improves by 20 percentage points, from 66 percent to 86 percent, and the effect is strongest on tasks that involve many contacts and two-handed cooperation.
  • The video-pretrained model reaches higher performance with only half the fine-tuning budget, which greatly improves training efficiency.
  • The same trend appears on real surgical robot data from JIGSAW, which confirms that the effect is real and not just a simulator artifact.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)