Latest AI and machine learning research in ophthalmology for healthcare professionals.
Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about images. Since its inception in 2015, VQA has rapidly evolved, driven by advances in deep learning, attention mechanisms, and transformer-based models. This survey traces the jou...
3D medical image segmentation has progressed considerably due to Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), yet these methods struggle to balance long-range dependency acquisition with computational efficiency. To address this challenge, we propose UNETVL (U-Net Vision-LSTM), a novel architecture that leverages recent advancements in temporal information processing. UNE...
Enhanced visual understanding serves as a cornerstone for multimodal large language models (MLLMs). Recent hybrid MLLMs incorporate a mixture of vis...
This paper introduces Shake-VLA, a Vision-Language-Action (VLA) model-based system designed to enable bimanual robotic manipulation for automated co...
Many documents, that we call templatized documents, are programmatically generated by populating fields in a visual template. Effective data extract...
We aim to assist image-based myopia screening by resolving two longstanding problems, "how to integrate the information of ocular images of a pair o...
Graph Neural Networks (GNNs) have become the standard approach for learning and reasoning over relational data, leveraging the message-passing mecha...
Current multimodal large language models (MLLMs) often underperform on mathematical problem-solving tasks that require fine-grained visual understan...
What types of numeric representations emerge in neural systems? What would a satisfying answer to this question look like? In this work, we interpre...
Transductive zero-shot learning with vision-language models leverages image-image similarities within the dataset to achieve better classification a...
The skin, as the largest organ of the human body, is vulnerable to a diverse array of conditions collectively known as skin lesions, which encompass...
Depicting novel classes with language descriptions by observing few-shot samples is inherent in human-learning systems. This lifelong learning capab...
Purpose: Diabetic retinopathy (DR) is a major cause of vision loss, particularly in India, where access to retina specialists is limited in rural ar...
Infants develop complex visual understanding rapidly, even preceding the acquisition of linguistic skills. As computer vision seeks to replicate the...
Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a l...
Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is ofte...
Recently, Visual Autoregressive ($\mathsf{VAR}$) Models introduced a groundbreaking advancement in the field of image generation, offering a scalabl...
While most existing neural image compression (NIC) and neural video compression (NVC) methodologies have achieved remarkable success, their optimiza...
Deep Anterior Lamellar Keratoplasty (DALK) is a partial-thickness corneal transplant procedure used to treat corneal stromal diseases. A crucial ste...
The emergence of artificial intelligence (AI), particularly deep learning (DL), has marked a new era in the realm of ophthalmology, offering transfo...