Latest AI and machine learning research in surgery for healthcare professionals.
Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches focus on coarse-grained, object-level localization, while traditional robotic grasping methods rely predominantly on geometric cues and lack language guidance, whi...
Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video MLLMs, trained primarily under a Supervised Fine-Tuning (SFT) paradigm, function as passive "Observers" that recognize ongoing events rather than evaluating the current state relative to the final task goal. In this paper, we introduce PRIMO R1 (Process Reason...
In aging-in-place contexts, small difficulties in Activities of Daily Living (ADL) can accumulate, affecting well-being through fatigue, anxiety, redu...
Accurate prediction of biochemical recurrence (BCR) after radical prostatectomy is critical for guiding adjuvant treatment and surveillance decisions ...
Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding...
Surgical scene understanding demands not only accurate predictions but also interpretable reasoning that surgeons can verify against clinical expertis...
Large language models (LLMs) increasingly guide clinical decisions through population-level evidence, yet they cannot encode individual patient prefer...
The unrestrained proliferation of cells that are malignant in nature is cancer. In recent times, medical professionals are constantly acquiring enhanc...
Introduction Postoperative complications after major surgery have substantial impacts on morbidity and resource utilisation. We investigated whether a...
A basic unanswered question in neural network training is: what is the best learning rate schedule shape for a given workload? The choice of learning ...
Mass spectrometry imaging (MSI) enables label-free visualization of molecular distributions across tissue samples but generates large and complex data...
Abstract Background: Echocardiography (echo) notes contain valuable prognostic information for patients in the intensive care unit (ICU). However, the...
Surgical scene Multi-Task Federated Learning (MTFL) is essential for robot-assisted minimally invasive surgery (RAS) but remains underexplored in surg...
Background: Pavlovian responding is a core component of behavior and can be measured via Pavlovian-instrumental transfer (PIT), where Pavlovian respon...
Accurate 3D reconstruction of vertebral anatomy from ultrasound is important for guiding minimally invasive spine interventions, but it remains challe...
Purpose: In this paper, we present a novel approach for online object tracking in laparoscopic cholecystectomy (LC) surgical videos, targeting localis...
Dynamic vision sensors, also known as event cameras, are rapidly rising in popularity for robotic and computer vision tasks due to their sparse activa...
Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on sta...
In the dynamic landscape of modern healthcare, maintaining the highest standards in surgical instruments is critical for clinical success. This report...
Rationale Autonomic dysfunction is a hallmark of sepsis pathophysiology, yet its quantification remains challenging. Multiscale entropy (MSE) derived ...