Latest AI and machine learning research in ophthalmology for healthcare professionals.
The Mary Tyler Moore Vision Initiative (MTM Vision) honors Mary Tyler Moore's commitment to ending vision loss from diabetes. Founded by Moore's husband, Dr. S. Robert Levine, MTM Vision aims to accelerate breakthroughs in diabetic retinal disease (DRD). At the MTM Vision Symposium 2024 on Curing Vision Loss from Diabetes, experts highlighted the urgent need for updated DRD staging systems, clinic...
Vision-and-Language Navigation (VLN) is a challenging task where an agent must understand language instructions and navigate unfamiliar environments using visual cues. The agent must accurately locate the target based on visual information from the environment and complete tasks through interaction with the surroundings. Despite significant advancements in this field, two major limitations persi...
Recent advancements in multimodal large language models (MLLMs) have broadened the scope of vision-language tasks, excelling in applications like im...
The rapid evolution of social media has provided enhanced communication channels for individuals to create online content, enabling them to express ...
The development of powerful user representations is a key factor in the success of recommender systems (RecSys). Online platforms employ a range of ...
Detecting plant diseases is a crucial aspect of modern agriculture, as it plays a key role in maintaining crop health and increasing overall yield. ...
Visual text is a crucial component in both document and scene images, conveying rich semantic information and attracting significant attention in th...
Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple o...
Diabetic retinopathy is a severe eye condition caused by diabetes where the retinal blood vessels get damaged and can lead to vision loss and blindn...
Facial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances in deep learni...
Identifying physiological and behavioral markers for mental health conditions is a longstanding challenge in psychiatry. Depression and suicidal ide...
Textual prompt tuning adapts Vision-Language Models (e.g., CLIP) in federated learning by tuning lightweight input tokens (or prompts) on local clie...
This work presents an RGB-D imaging-based approach to marker-free hand-eye calibration using a novel implementation of the iterative closest point (...
We propose a tightly-coupled LiDAR/Polarization Vision/Inertial/Magnetometer/Optical Flow Odometry via Smoothing and Mapping (LPVIMO-SAM) framework,...
Medical image reporting (MIR) aims to generate structured clinical descriptions from radiological images. Existing methods struggle with fine-graine...
By mapping sites at large scales using remotely sensed data, archaeologists can generate unique insights into long-term demographic trends, inter-re...
Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image in...
In this paper, we introduce AffectVLM, a vision-language model designed to integrate multiviews for a semantically rich and visually comprehensive u...
Graph Neural Networks (GNNs) have emerged as an efficient alternative to convolutional approaches for vision tasks such as image classification, lev...
Large Vision-Language Models (LVLMs) are pivotal for real-world AI tasks like embodied intelligence due to their strong vision-language reasoning ab...