Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4981-5000 of 9,853 articles

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

Medical ultrasonography is an essential imaging technique for examining superficial organs and tissues, including lymph nodes, breast, and thyroid. It employs high-frequency ultrasound waves to generate detailed images of the internal structures of the human body. However, manually contouring regions of interest in these images is a labor-intensive task that demands expertise and often results i...

ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction

Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still struggles with fine-grained, region-level understanding, hindering its effectiveness on these dense prediction tasks. We identify two pivotal factors required to address this limitation: semantic coherence and fine-gra...

Parallel FFTW on RISC-V: A Comparative Study including OpenMP, MPI, and HPX

Rapid advancements in RISC-V hardware development shift the focus from low-level optimizations to higher-level parallelization. Recent RISC-V proces...

Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers

Time series classification is a fundamental task in healthcare and industry, yet the development of time series foundation models (TSFMs) remains li...

Guidelines for Gaze-based Neural Preliminary Diagnosis

Neural disorders refer to any condition affecting the nervous system and that influence how individuals perceive and interact with the world. Tradit...

CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA

Vision Transformers (ViTs) exhibit superior performance in computer vision tasks but face deployment challenges on resource-constrained devices due ...

Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring

The application of visual instruction tuning and other post-training techniques has significantly enhanced the capabilities of Large Language Models...

SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding

Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a crit...

MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding

Different medical imaging modalities capture diagnostic information at varying spatial resolutions, from coarse global patterns to fine-grained loca...

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach ...

Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations

Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise...

Diffusion Counterfactual Generation with Semantic Abduction

Counterfactual image generation presents significant challenges, including preserving identity, maintaining perceptual quality, and ensuring faithfu...

Image Reconstruction as a Tool for Feature Analysis

Vision encoders are increasingly used in modern applications, from vision-only models to multimodal systems such as vision-language models. Despite ...

ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models

Architectural cultures across regions are characterized by stylistic diversity, shaped by historical, social, and technological contexts in addition...

Learning Speaker-Invariant Visual Features for Lipreading

Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract v...

APTOS-2024 challenge report: Generation of synthetic 3D OCT images from fundus photographs

Optical Coherence Tomography (OCT) provides high-resolution, 3D, and non-invasive visualization of retinal layers in vivo, serving as a critical too...

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests

Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics re...

MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologi...

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual diffe...

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Recent generations of language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answe...

Browse Categories