Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5141-5160 of 9,853 articles

Medical Large Vision Language Models with Multi-Image Visual Ability

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remains underexplored. Unlike single image based tasks, medical tasks involving multiple images often demand sophisticated visual understanding capabilities, such as temporal reasonin...

CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation

Most publicly available medical segmentation datasets are only partially labeled, with annotations provided for a subset of anatomical structures. When multiple datasets are combined for training, this incomplete annotation poses challenges, as it limits the model's ability to learn shared anatomical representations among datasets. Furthermore, vision-only frameworks often fail to capture comple...

Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering

Vision-Language Models (VLMs) have shown promise in various 2D visual tasks, yet their readiness for 3D clinical diagnosis remains unclear due to st...

Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

Large Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of ...

Eye-See-You: Reverse Pass-Through VR and Head Avatars

Virtual Reality (VR) headsets, while integral to the evolving digital ecosystem, present a critical challenge: the occlusion of users' eyes and port...

GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains

Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization pres...

EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries ...

Mind Your Vision: Multimodal Estimation of Refractive Disorders Using Electrooculography and Eye Tracking

Refractive errors are among the most common visual impairments globally, yet their diagnosis often relies on active user participation and clinical ...

How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control

We explore Large Language Models (LLMs)' human motion knowledge through 3D avatar control. Given a motion instruction, we prompt LLMs to first gener...

Cultural Awareness in Vision-Language Models: A Cross-Country Exploration

Vision-Language Models (VLMs) are increasingly deployed in diverse cultural contexts, yet their internal biases remain poorly understood. In this wo...

ImLPR: Image-based LiDAR Place Recognition using Vision Foundation Models

LiDAR Place Recognition (LPR) is a key component in robotic localization, enabling robots to align current scans with prior maps of their environmen...

COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification

We introduce the Convolutional Low-Rank Adaptation (CoLoRA) method, designed explicitly to overcome the inefficiencies found in current CNN fine-tun...

TokBench: Evaluating Your Visual Tokenizer before Visual Generation

In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate recon...

CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays

Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual qu...

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals

Large language models (LLMs) have demonstrated significant success in complex reasoning tasks such as math and coding. In contrast to these tasks wh...

Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling

Vision-language models (VLMs) have recently been integrated into multiple instance learning (MIL) frameworks to address the challenge of few-shot, w...

Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling

Vision-language models (VLMs) have recently been integrated into multiple instance learning (MIL) frameworks to address the challenge of few-shot, w...

AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models

Medical image segmentation is vital for clinical diagnosis, yet current deep learning methods often demand extensive expert effort, i.e., either thr...

KITINet: Kinetics Theory Inspired Network Architectures with PDE Simulation Approaches

Despite the widely recognized success of residual connections in modern neural networks, their design principles remain largely heuristic. This pape...

An Attention Infused Deep Learning System with Grad-CAM Visualization for Early Screening of Glaucoma

This research work reveals the eye opening wisdom of the hybrid labyrinthine deep learning models synergy born out of combining a trailblazing convo...

Browse Categories