Latest AI and machine learning research in ophthalmology for healthcare professionals.
The evolution of colour vision is captivating, as it reveals the adaptive strategies of extinct species while simultaneously inspiring innovations in modern imaging technology. In this study, we present a simplified model of visual transduction in the retina, introducing a novel opsin layer. We quantify evolutionary pressures by measuring machine vision recognition accuracy on colour images shap...
Recent advances in Large Vision-Language Models (LVLMs) have enabled general-purpose vision tasks through visual instruction tuning. While existing LVLMs can generate segmentation masks from text prompts for single images, they struggle with segmentation-grounded reasoning across images, especially at finer granularities such as object parts. In this paper, we introduce the new task of part-focu...
Current multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals though they give comprehensive perce...
With the rapid development of multimodal learning, the image-text matching task, as a bridge connecting vision and language, has become increasingly...
Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representatio...
3D Gaussian Splatting (3DGS) has gained significant attention for 3D scene reconstruction, but still suffers from complex outdoor environments, espe...
We present a novel method that extends the self-attention mechanism of a vision transformer (ViT) for more accurate object detection across diverse ...
Multi-head self-attention (MHSA) is a key component of Transformers, a widely popular architecture in both language and vision. Multiple heads intui...
Street view imagery (SVI) has been instrumental in many studies in the past decade to understand and characterize street features and the built envi...
Distinguishing spatial relations is a basic part of human cognition which requires fine-grained perception on cross-instance. Although benchmarks li...
While learned image compression methods have achieved impressive results in either human visual perception or machine vision tasks, they are often s...
Recent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. Howe...
Few-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Exi...
Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization...
The emergence of Vision-Language Models (VLMs) is a significant advancement in integrating computer vision with Large Language Models (LLMs) to enha...
Multimodal large language models (MLLMs) have significantly advanced tasks like caption generation and visual question answering by integrating visu...
Methods for converting the tacit knowledge of experts into explicit knowledge have drawn increasing attention. Gaze data has emerged as a valuable a...
Visual guidance (VG) is critical for directing user attention in virtual and augmented reality applications. However, conventional methods using exp...
Retinal image analysis is crucial for diagnosing and treating eye diseases, yet generating accurate medical reports from images remains challenging ...
The integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) opens new avenues for addressing complex challenges in multimodal ...