Table of Contents
Fetching ...

Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning

Lukas Zbinden, Nigel Nelson, Juo-Tung Chen, Xinhao Chen, Ji Woong Kim, Mahdi Azizian, Axel Krieger, Sean Huver

TL;DR

Cosmos-Surg-dVRK introduces a domain-specific world foundation model finetuned from Cosmos to simulate action-conditioned surgical rollouts and automatically evaluate policies via a video classifier. The approach demonstrates strong alignment with real-dVRK performance on tabletop suturing and shows promising results for ex-vivo cholecystectomy, highlighting the potential of learned simulators as cost-effective benchmarks. A V-JEPA 2-based classifier enables automated labeling that closely tracks human judgments, and incorporating failure trajectories during training improves realism and evaluation reliability. These findings suggest a pathway toward faster, safer, and more reproducible benchmarking of surgical autonomy policies, with broader implications for sim-to-real transfer and domain-specific digital twins in surgery.

Abstract

The rise of surgical robots and vision-language-action models has accelerated the development of autonomous surgical policies and efficient assessment strategies. However, evaluating these policies directly on physical robotic platforms such as the da Vinci Research Kit (dVRK) remains hindered by high costs, time demands, reproducibility challenges, and variability in execution. World foundation models (WFM) for physical AI offer a transformative approach to simulate complex real-world surgical tasks, such as soft tissue deformation, with high fidelity. This work introduces Cosmos-Surg-dVRK, a surgical finetune of the Cosmos WFM, which, together with a trained video classifier, enables fully automated online evaluation and benchmarking of surgical policies. We evaluate Cosmos-Surg-dVRK using two distinct surgical datasets. On tabletop suture pad tasks, the automated pipeline achieves strong correlation between online rollouts in Cosmos-Surg-dVRK and policy outcomes on the real dVRK Si platform, as well as good agreement between human labelers and the V-JEPA 2-derived video classifier. Additionally, preliminary experiments with ex-vivo porcine cholecystectomy tasks in Cosmos-Surg-dVRK demonstrate promising alignment with real-world evaluations, highlighting the platform's potential for more complex surgical procedures.

Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning

TL;DR

Cosmos-Surg-dVRK introduces a domain-specific world foundation model finetuned from Cosmos to simulate action-conditioned surgical rollouts and automatically evaluate policies via a video classifier. The approach demonstrates strong alignment with real-dVRK performance on tabletop suturing and shows promising results for ex-vivo cholecystectomy, highlighting the potential of learned simulators as cost-effective benchmarks. A V-JEPA 2-based classifier enables automated labeling that closely tracks human judgments, and incorporating failure trajectories during training improves realism and evaluation reliability. These findings suggest a pathway toward faster, safer, and more reproducible benchmarking of surgical autonomy policies, with broader implications for sim-to-real transfer and domain-specific digital twins in surgery.

Abstract

The rise of surgical robots and vision-language-action models has accelerated the development of autonomous surgical policies and efficient assessment strategies. However, evaluating these policies directly on physical robotic platforms such as the da Vinci Research Kit (dVRK) remains hindered by high costs, time demands, reproducibility challenges, and variability in execution. World foundation models (WFM) for physical AI offer a transformative approach to simulate complex real-world surgical tasks, such as soft tissue deformation, with high fidelity. This work introduces Cosmos-Surg-dVRK, a surgical finetune of the Cosmos WFM, which, together with a trained video classifier, enables fully automated online evaluation and benchmarking of surgical policies. We evaluate Cosmos-Surg-dVRK using two distinct surgical datasets. On tabletop suture pad tasks, the automated pipeline achieves strong correlation between online rollouts in Cosmos-Surg-dVRK and policy outcomes on the real dVRK Si platform, as well as good agreement between human labelers and the V-JEPA 2-derived video classifier. Additionally, preliminary experiments with ex-vivo porcine cholecystectomy tasks in Cosmos-Surg-dVRK demonstrate promising alignment with real-world evaluations, highlighting the platform's potential for more complex surgical procedures.
Paper Structure (31 sections, 13 figures, 5 tables)

This paper contains 31 sections, 13 figures, 5 tables.

Figures (13)

  • Figure 1: Online evaluation of surgical robot policies in Cosmos-Surg-dVRK simulation. We propose a framework for automated policy evaluation using Cosmos-Surg-dVRK, a Cosmos world foundation model (WFM) finetune, to perform simulated surgical policy rollouts and subsequent automated success rate evaluation using a video classifier. Evaluation proceeds by initializing the policy with an observed frame at state $s_0$. Conditioned on $s_0$, the policy and Cosmos-Surg-dVRK then generate $K$ action-conditioned future frames autoregressively, $s_{1:K}$. In each iteration, the predicted frames are appended to the output video, and the last predicted frame is used as input to the surgical policy for the next iteration. After rollout completion, the generated video is automatically labeled for task success or failure using a trained video classifier, enabling objective selection of the most promising surgical policies for real-robot evaluation and deployment.
  • Figure 2: Soft tissue simulation with Cosmos-Surg-dVRK. Surgical simulation in ex-vivo porcine cholecystectomy. Top row: example as recorded on the real dVRK Si. Middle row: Cosmos-Surg-dVRK generated example with identical kinematic action trajectory. Bottom row: Holdout L1 and SSIM vs. number of generated frames.
  • Figure 3: Surgical autonomy dataset distributions. a) Tabletop dataset consists of success episodes, recoveries, as well as failure data. The label needle pickup includes task needle handover. b) Ex-vivo porcine cholecystectomy dataset consists of success and recovery episodes.
  • Figure 4: Manual Cosmos-Surg-dVRK vs. real-world tabletop success rates. Relationship between manual surgical policy success rate evaluation in Cosmos-Surg-dVRK simulation (vertical axis) and on the real-world dVRK (horizontal axis). Each policy is shown with two training regimes: half–training and full–training across four tabletop suture pad tasks.
  • Figure 5: Qualitative examples of tabletop policy online rollouts in Cosmos-Surg-dVRK. Each row represents one different policy rollout. Top row: Successful example as recorded on the real dVRK. Middle two rows: Successfully completed tasks in Cosmos-Surg-dVRK. Bottom two rows: Examples of failed tasks in Cosmos-Surg-dVRK. Four frames per task are chosen individually to best represent the example.
  • ...and 8 more figures