Latest AI and machine learning research in ophthalmology for healthcare professionals.
Purpose: New therapeutic strategies such as optogenetics have created a need for accurate tracking of inner retina degeneration in Retinitis pigmentosa (RP) patients. We introduce two tailored deep learning models to segment the RNFL (retinal nerve fibre layer), GCIPL (ganglion cell inner plexiform layer), INL (inner nuclear layer), CFT (central foveal thickness) and RPE (retinal pigment epitheliu...
Objective. To preserve the encoding of visual information in prosthetic vision as close to natural as possible, subretinal photovoltaic implants, which replace the lost photoreceptors, strive to stimulate the second-order retinal neurons, the bipolar cells, while avoiding direct activation of the downstream retinal ganglion cells. To assess the range of such selective subretinal activation, we imp...
Vision-language models (VLMs) are rapidly advancing toward sophisticated grounded structured visual reasoning. Training models for such advanced capab...
Optical coherence tomography (OCT) is essential in ophthalmology, but inconsistent image quality especially in low-cost devices hinders automated anal...
Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception and reasoning capabilities. While numer...
Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-so...
Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, envi...
Objective: To systematically review automated nerve morphometry tools and independently benchmark their performance on independent optic nerve dataset...
Visual perception depends on top-down goals and bottom-up sensory mechanisms. Vision-language models implement both, allowing us to treat each compone...
Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, envi...
Visual information helps resolve ambiguity in coreference resolution, leading to notable performance gains. However, existing Multi-modal Coreference ...
When vision contradicts text, multimodal large language models (MLLMs) consistently favor text, even when images provide clear evidence otherwise. Thi...
Visual attribution is a fundamental tool for interpreting modern vision and vision-language models, particularly when their decisions must be inspecte...
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely ...
Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior wor...
Purpose: To evaluate the efficacy of large language models (LLMs) in extracting medication-related information from glaucoma clinical notes in the ele...
Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributional shif...
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for ...
Human perception of visual scenes is inherently temporal. We instinctively recognise whether a fruit is ripening or rotting, whether construction is p...
Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective predict...