Latest AI and machine learning research in ophthalmology for healthcare professionals.
Retinal foundation models have significantly advanced retinal image analysis by leveraging self-supervised learning to reduce dependence on labeled data while achieving strong generalization. Many recent approaches enhance retinal image understanding using report supervision, but obtaining clinical reports is often costly and challenging. In contrast, metadata (e.g., age, gender) is widely avail...
Optical Coherence Tomography (OCT) provides valuable insights in ophthalmology, cardiology, and neurology due to high-resolution, cross-sectional images of the retina. One critical task for ophthalmologists using OCT is delineation of retinal layers within scans. This process is time-consuming and prone to human bias, affecting the accuracy and reliability of diagnoses. Previous efforts to autom...
Although large Vision-Language Models (VLMs) have demonstrated remarkable performance in a wide range of multimodal tasks, their true reasoning capa...
Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-par...
Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domain...
For the elderly population, falls pose a serious and increasing risk of serious injury and loss of independence. In order to overcome this difficult...
Current research has explored vision-language models for multi-modal embedding tasks, such as information retrieval, visual grounding, and classific...
Document retrieval is an important task for search and Retrieval-Augmented Generation (RAG) applications. Large Language Models (LLMs) have contribu...
Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across div...
To perform autonomous visual search for environmental monitoring, a robot may leverage satellite imagery as a prior map. This can help inform coarse...
Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-t...
Vision-Language-Action (VLA) models have recently become highly prominent in the field of robotics. Leveraging vision-language foundation models tra...
Recent advances in self-supervised learning (SSL) have revolutionized computer vision through innovative architectures and learning objectives, yet ...
Recent advances in diffusion models have demonstrated their remarkable ability to capture complex image distributions, but the geometric properties ...
Accurate segmentation of regions of interest in biomedical images holds substantial value in image analysis. Although several foundation models for ...
Minimally invasive surgery (MIS) presents significant visual and technical challenges, including surgical instrument classification and understandin...
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are cap...
Existing vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can gene...
Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical ...
PURPOSE: To evaluate the accuracy of 11 intraocular lens (IOL) calculation formulas in eyes undergoing Descemet membrane endothelial keratoplasty (DME...