Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4441-4460 of 9,853 articles

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and significantly raises computational costs. We attribute this inefficiency to substantial visual regions r...

Mar 25 2026 2603.24326v1

Vision-based Deep Learning Analysis of Unordered Biomedical Tabular Datasets via Optimal Spatial Cartography

Tabular data are central to biomedical research, from liquid biopsy and bulk and single-cell transcriptomics to electronic health records and phenotypic profiling. Unlike images or sequences, however, tabular datasets lack intrinsic spatial organization: features are treated as unordered dimensions, and their relationships must be inferred implicitly by the model. This limits the ability of vision...

Mar 24 2026 2603.22675v1
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large L...

Mar 24 2026 2603.22815v1
Curriculum-Driven 3D CT Report Generation via Language-Free Visual Grafting and Zone-Constrained Compression

Automated radiology report generation from 3D computed tomography (CT) volumes is challenging due to extreme sequence lengths, severe class imbalance,...

Mar 24 2026 2603.23308v1
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation

Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming vid...

Mar 24 2026 2603.23491v1
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. T...

Mar 24 2026 2603.23495v1
Language Models Can Explain Visual Features via Steering

Sparse Autoencoders uncover thousands of features in vision models, yet explaining these features without requiring human intervention remains an open...

Mar 23 2026 2603.22593v1
Q-Tacit: Image Quality Assessment via Latent Visual Reasoning

Vision-Language Model (VLM)-based image quality assessment (IQA) has been significantly advanced by incorporating Chain-of-Thought (CoT) reasoning. Re...

Mar 23 2026 2603.22641v1
Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models

Continual unlearning poses the challenge of enabling large vision-language models to selectively refuse specific image-instruction pairs in response t...

Mar 23 2026 2603.21484v1
CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation

We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract opht...

Mar 23 2026 2603.21566v1
Rethinking Token Reduction for Large Vision-Language Models

Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Althou...

Mar 23 2026 2603.21701v1
SteelDefectX: A Coarse-to-Fine Vision-Language Dataset and Benchmark for Generalizable Steel Surface Defect Detection

Steel surface defect detection is essential for ensuring product quality and reliability in modern manufacturing. Current methods often rely on basic ...

Mar 23 2026 2603.21824v1
6D Robotic OCT Scanning of Curved Tissue Surfaces

Optical coherence tomography (OCT) is a non-invasive volumetric imaging modality with high spatial and temporal resolution. For imaging larger tissue ...

Mar 23 2026 2603.22012v1
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning

Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are pr...

Mar 23 2026 2603.22057v1
P-Flow: Prompting Visual Effects Generation

Recent advancements in video generation models have significantly improved their ability to follow text prompts. However, the customization of dynamic...

Mar 23 2026 2603.22091v1
The Dual Mechanisms of Spatial Reasoning in Vision-Language Models

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to associate objects with their p...

Mar 23 2026 2603.22278v1
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented ...

Mar 22 2026 2603.21165v1
CornOrb: A Multimodal Dataset of Orbscan Corneal Topography and Clinical Annotations for Keratoconus Detection

In this paper, we present CornOrb, a publicly accessible multimodal dataset of Orbscan corneal topography images and clinical annotations collected fr...

Mar 22 2026 2603.21245v1
Image-Based Structural Analysis Using Computer Vision and LLMs: PhotoBeamSolver

This paper presents the development of a documented program capable of solving idealized beam models, such as those commonly used in textbooks and aca...

Mar 22 2026 2603.21432v1
SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification

Monocular depth estimation remains challenging for transparent objects, where refraction and transmission are difficult to model and break the appeara...

Mar 20 2026 2603.19547v1
Browse Categories