Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4541-4560 of 9,853 articles

POLISH'ing the Sky: Wide-Field and High-Dynamic Range Interferometric Image Reconstruction with Application to Strong Lens Discovery

Radio interferometry enables high-resolution imaging of astronomical radio sources by synthesizing a large effective aperture from an array of antennas and solving a deconvolution problem to reconstruct the image. Deep learning has emerged as a promising solution to the imaging problem, reducing computational costs and enabling super-resolution. However, existing DL-based methods often fall short ...

Mar 10 2026 2603.09162v1

MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data

Self-evolving has emerged as a key paradigm for improving foundational models such as Large Language Models (LLMs) and Vision Language Models (VLMs) with minimal human intervention. While recent approaches have demonstrated that LLM agents can self-evolve from scratch with little to no data, VLMs introduce an additional visual modality that typically requires at least some seed data, such as image...

Mar 10 2026 2603.09206v1
OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision language tasks. However, their ability in ...

Mar 10 2026 2603.09326v1
OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks

Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest i...

Mar 10 2026 2603.09471v1
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

Vision-language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows th...

Mar 10 2026 2603.09480v1
A comprehensive study of time-of-flight non-line-of-sight imaging

Time-of-Flight non-line-of-sight (ToF NLOS) imaging techniques provide state-of-the-art reconstructions of scenes hidden around corners by inverting t...

Mar 10 2026 2603.09548v1
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning

Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approach...

Mar 10 2026 2603.09566v1
A saccade-inspired approach to image classification using visiontransformer attention maps

Human vision achieves remarkable perceptual performance while operating under strict metabolic constraints. A key ingredient is the selective attentio...

Mar 10 2026 2603.09613v1
Grounding Synthetic Data Generation With Vision and Language Models

Deep learning models benefit from increasing data diversity and volume, motivating synthetic data augmentation to improve existing datasets. However, ...

Mar 10 2026 2603.09625v1
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA ...

Mar 10 2026 2603.09689v1
VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecti...

Mar 10 2026 2603.09826v1
BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion

Language-conditioned local navigation requires a robot to infer a nearby traversable target location from its current observation and an open-vocabula...

Mar 10 2026 2603.09961v1
Computer Vision-Based Vehicle Allotment System using Perspective Mapping

Smart city research envisions a future in which data-driven solutions and sustainable infrastructure work together to define urban living at the cross...

Mar 9 2026 2603.08827v1
A Tutorial on Automated Classification of Eye Diseases Using Deep Learning

Sight is one of the five senses essential to human experience, and the eyes are vital organs that require careful protection. These organs are also su...

Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision-Language Models

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive multimodal understanding in general domains. However, their applicabilit...

Mar 9 2026 2603.07868v1
Text to Automata Diagrams: Comparing TikZ Code Generation with Direct Image Synthesis

Diagrams are widely used in teaching computer science courses. They are useful in subjects such as automata and formal languages, data structures, etc...

Mar 9 2026 2603.07936v1
VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer

Zero-shot anomaly detection (ZSAD) requires detecting and localizing anomalies without access to target-class anomaly samples. Mainstream methods rely...

Mar 9 2026 2603.07952v1
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation

Existing aerial Vision-Language Navigation (VLN) methods predominantly adopt a detection-and-planning pipeline, which converts open-vocabulary detecti...

Mar 9 2026 2603.08007v1
See and Switch: Vision-Based Branching for Interactive Robot-Skill Programming

Programming robots by demonstration (PbD) is an intuitive concept, but scaling it to real-world variability remains a challenge for most current teach...

Mar 9 2026 2603.08057v1
Exploring Deep Learning and Ultra-Widefield Imaging for Diabetic Retinopathy and Macular Edema

Diabetic retinopathy (DR) and diabetic macular edema (DME) are leading causes of preventable blindness among working-age adults. Traditional approache...

Mar 9 2026 2603.08235v1
Browse Categories