TempoSafe-CVS: A Temporal Multi-scale Deep Learning Framework for Automated Assessment of the Critical View of Safety in Laparoscopic Cholecystectomy.
Journal:
Journal of imaging informatics in medicine
Published Date:
Aug 10, 2026
Abstract
Accurate assessment of the critical view of safety (CVS) is essential for preventing bile duct injuries during laparoscopic cholecystectomy. Existing artificial intelligence approaches primarily rely on static frame-level analysis and often fail to capture the temporal evolution of surgical scenes, limiting their ability to provide reliable and context-aware safety assessment. To address this challenge, we propose TempoSafe-CVS, a temporal multi-scale framework for automated CVS assessment in surgical videos. The proposed architecture integrates complementary visual representations through a Swin Transformer-based global context encoder, a ResNet-based local feature extractor, and a structure-aware convolutional module. These multi-scale features are combined and processed using temporal sequence modelling and spatio-temporal reasoning to capture both visual and temporal dependencies across surgical sequences. Furthermore, a unified multi-task prediction framework jointly estimates CVS safety status, procedural progression, anatomical structure visibility, and clinically relevant C1/C2/C3 criteria. Experiments conducted on the Endoscapes benchmark dataset demonstrate the effectiveness of the proposed approach, achieving 79.6% AUC-ROC for safety assessment, 81.5% average balanced accuracy for C1/C2/C3 criteria classification, and a mean absolute error of 0.187 for progression estimation. Comparative evaluations show consistent improvements over existing CVS assessment methods, highlighting the benefits of temporal reasoning and multi-scale visual representation learning. Qualitative analyses further demonstrate the interpretability of the framework through temporally consistent and anatomically grounded predictions. The proposed framework advances intelligent surgical video understanding by combining temporal sequence reasoning with multi-scale visual analysis, offering a potential solution for explainable and context-aware decision support in safety-critical surgical environments.
Authors
Keywords
No keywords available for this article.