Latest AI and machine learning research in ophthalmology for healthcare professionals.
Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphical perception tasks, which are essential for interpreting visualizations, the perceptual capabilities of ViTs remain largely unexplored. In this work, we investigate the performance...
Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are evident in a wide range of VLMs built on a variety of base models, alignment architectures, and training data. However, recent works show that these models trail behind in traditi...
Artificial intelligence (AI) has increasingly transformed medical prognostics by enabling rapid and accurate analysis across imaging and pathology. Ho...
Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying on...
Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. This la...
We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ...
Eyeframe lens tracing is an important process in the optical industry that requires sub-millimeter precision to ensure proper lens fitting and optimal...
Time-series anomaly detection (TSAD) requires identifying both immediate Point Anomalies and long-range Context Anomalies. However, existing foundatio...
Compositional generalization, the ability to reason about novel combinations of familiar concepts, is fundamental to human cognition and a critical ch...
Visual loco-manipulation of arbitrary objects in the wild with humanoid robots requires accurate end-effector (EE) control and a generalizable underst...
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data....
Visual hallucinations (VHs) occur across psychedelic states and diverse psychiatric and neurological conditions, yet their phenomenology remains diffi...
Type 1 narcolepsy (NT1), a disorder caused by the loss of hypocretin/orexin transmission, is characterized by daytime sleepiness and symptoms where Ra...
The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). Th...
Vision Language Models (VLMs) are designed to extend Large Language Models (LLMs) with visual capabilities, yet in this work we observe a surprising p...
Biological age estimators quantify aging-related variation but provide limited insight into organ-specific aging processes. The retina enables non-inv...
We address fine-grained visual reasoning in multimodal large language models (MLLMs), where key evidence may reside in tiny objects, cluttered regions...
Referring Image Segmentation (RIS) aims to segment a target object described by a natural language expression. Existing methods have evolved by levera...
Vision language models (VLMs) achieve strong performance on RGB imagery, but they do not generalize to thermal images. Thermal sensing plays a critica...
Data-driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most exi...