Large language model-augmented implicit surgical video review

Journal: medRxiv
Published Date:

Abstract

Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.

Authors

  • Zhang
  • Z.; Qadir
  • M. I.; Ramchand
  • R.; Belwadi
  • M.; Ball
  • R. P.; Konstantinopoulos
  • K.; Abbey
  • E. M.; Ernsberger
  • K. T.; Guzman
  • M. J.; Hendren
  • S.; Holcomb
  • B. K.; Robb
  • B. W.; Stankowski
  • T.; Waters
  • J. A.; Stefanidis
  • D.; Bilimoria
  • K. Y.; Mohanty
  • S.; Kolbinger
  • F. R.

Categories