Multimodal AI for automated assessment of surgical performance
Abstract
Accurate and objective assessment of surgical skill is crucial for improving training outcomes, ensuring patient safety, and enabling scalable evaluation across clinical and simulation contexts. Current approaches often depend on subjective human ratings or unimodal analysis, limiting reliability and applicability in real-world settings. We propose a novel multimodal deep learning framework that integrates visual and kinematic data to evaluate surgical proficiency automatically. Our method leverages YOLOv8 and BYTE Tracker for surgical tool tracking and trajectory extraction, an autoencoder for encoding tool motion, ResNet50 for deriving semantic features from heatmap density plots, and a SlowFast-R50 backbone for capturing spatio-temporal video cues. These latent features are fused and input to a 1D-convolutional neural network to predict expert-assigned skill scores. On a JIGSAWS simulation benchmark (consisting of 103 videos), our model achieved a Spearman correlation of 0.85 in 4-fold cross-validation and 0.87 in a leave-one-user-out evaluation, which outperforms prior approaches. On a Biotissue dataset (205 videos including GJ, PJ, HJ), fusing kinematic and visual features yielded a top Spearman correlation of 0.67 (GJ dataset) and 0.69 (PJ dataset). These results demonstrate our framework’s scalability, robustness, and potential for clinical deployment.
// Source
Authors: S. M. Khairnar, Huu Phong Nguyen, Sofia Garces-Palacios, SAMY CASTILLO, Andres Abreu, Amr Al Abbas, Kaustubh Gopal, Daniel Scott, Herbert J. Zeh, Patricio Polanco, Ganesh Sankaranarayanan
Institutions: Southwestern Medical Center, The University of Texas Southwestern Medical Center, Southwestern Medical Center