Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4581-4600 of 9,853 articles

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios. However, existing methods typically feed LLMs with tokens from multi-view and multi-frame images independently, leading to redundant computation and l...

Mar 6 2026 2603.06576v1

A Quantum Lens on Molecular Design: A Machine-Learned Energy Function from Interacting Quantum Atoms.

Accurate predictions of the interactions (covalent bonds and non-covalent contacts between atoms) in a molecular system require scalable, accurate, and interpretable energy functions. While classical force fields and knowledge-based energy functions struggle to capture key electronic effects, quantum chemistry approaches such as density functional theory (DFT) provide the necessary accuracy but re...

Locality-Attending Vision Transformer

Vision transformers have demonstrated remarkable success in classification by leveraging global self-attention to capture long-range dependencies. How...

Mar 5 2026 2603.04892v1
Location-Aware Pretraining for Medical Difference Visual Question Answering

Unlike conventional single-image models, differential medical VQA frameworks process multiple images to identify differences, mirroring the comparativ...

Mar 5 2026 2603.04950v1
Physics-consistent deep learning for blind aberration recovery in mobile optics

Mobile photography is often limited by complex, lens-specific optical aberrations. While recent deep learning methods approach this as an end-to-end d...

Mar 5 2026 2603.04999v1
Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model

Source Free Unsupervised Domain Adaptation (SFUDA) is critical for deploying deep learning models across diverse clinical settings. However, existing ...

Mar 5 2026 2603.05012v1
MI-DETR: A Strong Baseline for Moving Infrared Small Target Detection with Bio-Inspired Motion Integration

Infrared small target detection (ISTD) is challenging because tiny, low-contrast targets are easily obscured by complex and dynamic backgrounds. Conve...

Mar 5 2026 2603.05071v1
Augmenting representations with scientific papers

Astronomers have acquired vast repositories of multimodal data, including images, spectra, and time series, complemented by decades of literature that...

Mar 4 2026 2603.04516v1
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs

Multimodal Large Language Models (MLLMs) have shown strong performance in vision-language tasks, but their inference efficiency is severely limited by...

Mar 4 2026 2603.03681v1
Glass Segmentation with Fusion of Learned and General Visual Features

Glass surface segmentation from RGB images is a challenging task, since glass as a transparent material distinctly lacks visual characteristics. Howev...

Mar 4 2026 2603.03718v1
Separators in Enhancing Autoregressive Pretraining for Vision Mamba

The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting significant attention due to its efficient pro...

Mar 4 2026 2603.03806v1
When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models

When visual evidence is ambiguous, vision models must decide whether to interpret face-like patterns as meaningful. Face pareidolia, the perception of...

Mar 4 2026 2603.03989v1
Weakly Supervised Patch Annotation for Improved Screening of Diabetic Retinopathy

Diabetic Retinopathy (DR) requires timely screening to prevent irreversible vision loss. However, its early detection remains a significant challenge ...

Mar 4 2026 2603.03991v1
Beyond Mixtures and Products for Ensemble Aggregation: A Likelihood Perspective on Generalized Means

Density aggregation is a central problem in machine learning, for instance when combining predictions from a Deep Ensemble. The choice of aggregation ...

Mar 4 2026 2603.04204v1
Modeling Cross-vision Synergy for Unified Large Vision Model

Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, vi...

Mar 3 2026 2603.03564v1
CausalFund: Causality-Inspired Domain Generalization in Retinal Fundus Imaging for Low-Resource Screening

Early screening for glaucoma and diabetic retinopathy (DR) is critical to prevent irreversible vision loss, yet remains inaccessible to many underserv...

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques ...

Mar 3 2026 2603.02556v1
Beyond Language Modeling: An Exploration of Multimodal Pretraining

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space ...

Mar 3 2026 2603.03276v1
Estimating Visual Attribute Effects in Advertising from Observational Data: A Deepfake-Informed Double Machine Learning Approach

Digital advertising increasingly relies on visual content, yet marketers lack rigorous methods for understanding how specific visual attributes causal...

Mar 2 2026 2603.02359v1
Predicting visual function before glaucoma onset from baseline optical coherence tomography scans using deep learning

Background: The visual field (VF) test results of many eyes with glaucoma progress despite treatment. This suggests that some eyes are either untreate...

Browse Categories