Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4461-4480 of 9,853 articles

Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification

Lifelong person re-identification (LReID) aims to learn from varying domains to obtain a unified person retrieval model. Existing LReID approaches typically focus on learning from scratch or a visual classification-pretrained model, while the Vision-Language Model (VLM) has shown generalizable knowledge in a variety of tasks. Although existing methods can be directly adapted to the VLM, since they...

Mar 20 2026 2603.19678v1

Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy

Deep learning underlies most modern approaches and tools in computer vision, including biomedical imaging. However, for interactive semantic segmentation (often called pixel classification in this context) and interactive object-level classification (object classification), feature-based shallow learning remains widely used. This is due to the diversity of data in this domain, the lack of large pr...

Mar 20 2026 2603.19802v1
Real-Time Structural Detection for Indoor Navigation from 3D LiDAR Using Bird's-Eye-View Images

Efficient structural perception is essential for mapping and autonomous navigation on resource-constrained robots. Existing 3D methods are computation...

Mar 20 2026 2603.19830v1
MFil-Mamba: Multi-Filter Scanning for Spatial Redundancy-Aware Visual State Space Models

State Space Models (SSMs), especially recent Mamba architecture, have achieved remarkable success in sequence modeling tasks. However, extending SSMs ...

Mar 20 2026 2603.20074v1
Tinted Frames: Question Framing Blinds Vision-Language Models

Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In th...

Mar 19 2026 2603.19203v2
Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following

Large vision language models (LVLMs) have demonstrated impressive performance across a wide range of tasks. These capabilities largely stem from visua...

Mar 19 2026 2603.19482v1
Vision Tiny Recursion Model (ViTRM): Parameter-Efficient Image Classification via Recursive State Refinement

The success of deep learning in computer vision has been driven by models of increasing scale, from deep Convolutional Neural Networks (CNN) to large ...

Mar 19 2026 2603.19503v1
Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their applicatio...

Mar 19 2026 2603.19516v1
A Cross-Study Multi-Organ Cell Atlas ofMacaca fascicularis Informed by Human Foundation Model Annotation: A Resource for Translational Target Assessment

Non-human primates (NHPs), particularly Macaca fascicularis (cynomolgus macaque), represent an essential model for preclinical assessment of biologics...

Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models

Counting serves as a simple but powerful test of a Large Vision-Language Model's (LVLM's) reasoning; it forces the model to identify each individual o...

Mar 19 2026 2603.18523v1
CoDA: Exploring Chain-of-Distribution Attacks and Post-Hoc Token-Space Repair for Medical Vision-Language Models

Medical vision--language models (MVLMs) are increasingly used as perceptual backbones in radiology pipelines and as the visual front end of multimodal...

Mar 19 2026 2603.18545v1
SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation

Realistic and efficient 3D garment generation remains a longstanding challenge in computer vision and digital fashion. Existing methods typically rely...

Mar 19 2026 2603.19053v1
Tinted Frames: Question Framing Blinds Vision-Language Models

Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In th...

Mar 19 2026 2603.19203v1
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders

Large vision--language models (VLMs) often use a frozen vision backbone, whose image features are mapped into a large language model through a lightwe...

Mar 19 2026 2603.19209v1
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial...

Mar 19 2026 2603.19219v1
LLM-Augmented Computational Phenotyping of Long Covid

Phenotypic characterization is essential for understanding heterogeneity in chronic diseases and for guiding personalized interventions. Long COVID, a...

Mar 18 2026 2603.18115v1
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning

Visual-Language Models (VLMs) have achieved remarkable progress in image captioning, visual question answering, and visual reasoning. Yet they remain ...

Mar 18 2026 2603.18282v1
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs

Multimodal Large Language Models (MLLMs) are increasingly applied to pixel-level vision tasks, yet their intrinsic capacity for spatial understanding ...

Mar 18 2026 2603.17228v1
MedSAD-CLIP: Supervised CLIP with Token-Patch Cross-Attention for Medical Anomaly Detection and Segmentation

Medical anomaly detection (MAD) and segmentation play a critical role in assisting clinical diagnosis by identifying abnormal regions in medical image...

Mar 18 2026 2603.17325v1
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions

While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. ...

Mar 18 2026 2603.17326v1
Browse Categories