Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4761-4780 of 9,853 articles

Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions

Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, even though the visual change is obvious to humans. This raises a fundamental question: do VLMs perceive visual changes or merely recall memorized patterns? While several studies have noted this phenomenon, the underlying ...

Jan 29 2026 2601.22150v1

Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models

Recent research has shown that aligning fine-grained text descriptions with localized image patches can significantly improve the zero-shot performance of pre-trained vision-language models (e.g., CLIP). However, we find that both fine-grained text descriptions and localized image patches often contain redundant information, making text-visual alignment less effective. In this paper, we tackle thi...

Jan 28 2026 2601.20419v1
DeepSeek-OCR 2: Visual Causal Flow

We present DeepSeek-OCR 2 to investigate the feasibility of a novel encoder-DeepEncoder V2-capable of dynamically reordering visual tokens upon image ...

Jan 28 2026 2601.20552v1
CLEAR-Mamba:Towards Accurate, Adaptive and Trustworthy Multi-Sequence Ophthalmic Angiography Classification

Medical image classification is a core task in computer-aided diagnosis (CAD), playing a pivotal role in early disease detection, treatment planning, ...

Jan 28 2026 2601.20601v1
bi-modal textual prompt learning for vision-language models in remote sensing

Prompt learning (PL) has emerged as an effective strategy to adapt vision-language models (VLMs), such as CLIP, for downstream tasks under limited sup...

Jan 28 2026 2601.20675v1
Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification

"Compression Tells Intelligence", is supported by research in artificial intelligence, particularly concerning (multimodal) large language models (LLM...

Jan 28 2026 2601.20742v1
ProMist-5K: A Comprehensive Dataset for Digital Emulation of Cinematic Pro-Mist Filter Effects

Pro-Mist filters are widely used in cinematography for their ability to create soft halation, lower contrast, and produce a distinctive, atmospheric s...

Jan 27 2026 2601.19295v1
Robust Uncertainty Estimation under Distribution Shift via Difference Reconstruction

Estimating uncertainty in deep learning models is critical for reliable decision-making in high-stakes applications such as medical imaging. Prior res...

Jan 27 2026 2601.19341v1
PaW-ViT: A Patch-based Warping Vision Transformer for Robust Ear Verification

The rectangular tokens common to vision transformer methods for visual recognition can strongly affect performance of these methods due to incorporati...

Jan 27 2026 2601.19771v1
Fair-Eye Net: A Fair, Trustworthy, Multimodal Integrated Glaucoma Full Chain AI System

Glaucoma is a top cause of irreversible blindness globally, making early detection and longitudinal follow-up pivotal to preventing permanent vision l...

Jan 26 2026 2601.18464v1
A Computational Approach to Visual Metonymy

Images often communicate more than they literally depict: a set of tools can suggest an occupation and a cultural artifact can suggest a tradition. Th...

Jan 25 2026 2601.17706v1
ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative Pruning

Large Vision-Language Models (LVLMs) incur high computational costs due to significant redundancy in their visual tokens. To effectively reduce this c...

Jan 25 2026 2601.17818v1
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models

Large Vision-Language Models (LVLMs) hold significant promise for medical applications, yet their deployment is often constrained by insufficient alig...

Jan 25 2026 2601.17918v1
Population-Scale Analysis of Frequency-Dependent Calcium Dynamics in Retinal Ganglion Cells Under Electric Field Stimulation

Electric field (EF) stimulation is an emerging neuromodulatory strategy for promoting the repair and functional recovery of degenerated neural network...

SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision

Retinal prostheses restore limited visual perception, but low spatial resolution and temporal persistence make reading difficult. In sequential letter...

Jan 24 2026 2601.17326v1
Beyond the Focus of Expansion: Retinal curl as a functional signal for heading estimation

Prevailing models aiming at explaining heading assume that humans need to recover the Focus of Expansion (FOE) while filtering out rotational flow (cu...

VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly Detection

Few-Shot Anomaly Detection (FSAD) has emerged as a critical paradigm for identifying irregularities using scarce normal references. While recent metho...

Jan 23 2026 2601.16381v1
X-Aligner: Composed Visual Retrieval without the Bells and Whistles

Composed Video Retrieval (CoVR) facilitates video retrieval by combining visual and textual queries. However, existing CoVR frameworks typically fuse ...

Jan 23 2026 2601.16582v1
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

While text-to-image (T2I) models have advanced considerably, their capability to associate colors with implicit concepts remains underexplored. To add...

Jan 23 2026 2601.16836v1
Analyzing Neural Network Information Flow Using Differential Geometry

This paper provides a fresh view of the neural network (NN) data flow problem, i.e., identifying the NN connections that are most important for the pe...

Jan 22 2026 2601.16366v1
Browse Categories