Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3801-3820 of 9,853 articles

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leadin...

Aug 11 2026 2608.11167v1

Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach

This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned contact lenses has been extensively studied under the presentation attack detection (PAD) paradigm, clear prescription contact lenses, that are typically transparent, have received comparatively less attention despite their widespread use. Unlike pa...

Aug 10 2026 2608.08977v1
Learning human joint torques from pixels

Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-worl...

Aug 10 2026 2608.09083v1
Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition

Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR...

Aug 10 2026 2608.09100v1
Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize ...

Aug 10 2026 2608.09176v1
RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures ar...

Aug 10 2026 2608.09226v1
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views

Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and...

Aug 10 2026 2608.09270v1
Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assis...

Aug 10 2026 2608.09302v1
Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language ...

Aug 10 2026 2608.09344v1
Disentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert Routing

Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation...

Aug 10 2026 2608.09752v1
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, t...

Aug 9 2026 2608.08676v1
Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial...

Aug 9 2026 2608.08713v1
UPolarSQ: Polar Representation Learning for Optic Disc and Peripapillary Atrophy Segmentation and Quantification in Fundus Photographs

Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provi...

Aug 9 2026 2608.08771v1
SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport

Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific b...

Aug 9 2026 2608.08840v1
Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Mode...

Aug 9 2026 2608.08964v1
Deep Learning of Fluorescence Lifetime Imaging Ophthalmoscopy for Type 2 Diabetes Classification

Purpose: To evaluate whether fluorescence lifetime imaging ophthalmoscopy (FLIO) combined with deep learning can detect metabolic signatures for class...

Invasion status stratifies the composition of the human pancreatic cancer perineural niche

The peripheral nervous system innervates the pancreatic ductal adenocarcinoma (PDAC) microenvironment, and perineural invasion (PNI), the invasion of ...

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertai...

Aug 6 2026 2608.05683v1
Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approach...

Aug 6 2026 2608.05803v1
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case

Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural...

Aug 6 2026 2608.06075v1
Browse Categories