Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5561-5580 of 9,853 articles

GazeLLM: Multimodal LLMs incorporating Human Visual Attention

Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human activities through video and audio, enabling many human-computer interaction and human-augmentation applications such as human activity support, real-world agents, and skill tra...

Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation

We propose a novel approach that adapts hierarchical vision foundation models for real-time ultrasound image segmentation. Existing ultrasound segmentation methods often struggle with adaptability to new tasks, relying on costly manual annotations, while real-time approaches generally fail to match state-of-the-art performance. To overcome these limitations, we introduce an adaptive framework th...

CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success in cross-modal tasks such as zero-shot image classification and text-i...

SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation

Large vision and language models show strong performance in tasks like image captioning, visual question answering, and retrieval. However, challeng...

It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data

The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In ...

Exploring Temporal Dynamics in Event-based Eye Tracker

Eye-tracking is a vital technology for human-computer interaction, especially in wearable devices such as AR, VR, and XR. The realization of high-sp...

CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation

Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly either assess text-...

BiPVL-Seg: Bidirectional Progressive Vision-Language Fusion with Global-Local Alignment for Medical Image Segmentation

Medical image segmentation typically relies solely on visual data, overlooking the rich textual information clinicians use for diagnosis. Vision-lan...

Re-Aligning Language to Visual Objects with an Agentic Workflow

Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve L...

Evolutionary Prompt Optimization Discovers Emergent Multimodal Reasoning Strategies in Vision-Language Models

We present a framework for optimizing prompts in vision-language models to elicit multimodal reasoning without model retraining. Using an evolutiona...

CADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation

Referring Remote Sensing Image Segmentation (RRSIS) is a challenging task, aiming to segment specific target objects in remote sensing (RS) images b...

LaViC: Adapting Large Vision-Language Models to Visually-Aware Conversational Recommendation

Conversational recommender systems engage users in dialogues to refine their needs and provide more personalized suggestions. Although textual infor...

LSNet: See Large, Focus Small

Vision network designs, including Convolutional Neural Networks and Vision Transformers, have significantly advanced the field of computer vision. Y...

RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning

Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessi...

Evaluating Compositional Scene Understanding in Multimodal Generative Models

The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential...

Optimizing Library Usage and Browser Experience: Application to the New York Public Library

We tackle the challenge brought to urban library systems by the {holds system} -- which allows users to request books available at other branches to...

VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Large Vision-Language Models (LVLMs) struggle with puzzles, which require precise perception, rule comprehension, and logical reasoning. Assessing a...

Shape and Texture Recognition in Large Vision-Language Models

Shape and texture recognition is fundamental to visual perception. The ability to identify shapes regardless of orientation, texture, or context, an...

A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of...

Enhance Vision-based Tactile Sensors via Dynamic Illumination and Image Fusion

Vision-based tactile sensors use structured light to measure deformation in their elastomeric interface. Until now, vision-based tactile sensors suc...

Browse Categories