Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3741-3760 of 9,853 articles

WISE-Screen: A Smartphone-Based Analytical Framework for Automated ASD Screening and Phenotyping via High-Fidelity Eye-tracking

The rising prevalence of autism spectrum disorder (ASD) strains clinical infrastructure. Gold-standard tools like ADOS-2 face high costs, specialized training requirements, and extensive waitlists, delaying diagnosis and intervention. While eye-tracking offers a promising digital biomarker, existing tools lack scalable community deployment due to hardware costs and operational constraints. Here, w...

Hyperbolic Hierarchical Clustering for Visual Representation Learning

We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental component of modern vision backbones like vision Transformers, facilitating information exchange between image patches. Mainstream token mixers, which rely on convolution, attention, MLP, or their hybrids, primarily focus on ...

Aug 24 2026 2608.22665v1
ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution ...

Aug 24 2026 2608.22996v1
E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Exis...

Aug 24 2026 2608.23253v1
What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal ...

Aug 24 2026 2608.23474v1
EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete cras...

Aug 24 2026 2608.23563v1
When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understandin...

Aug 23 2026 2608.22174v1
AI supported in silico screening of chimeric antigen receptor therapy targets

Chimeric antigen receptor (CAR) cell therapy has achieved transformative clinical success through targeting of CD19 in refractory B cell malignancies,...

MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model

Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods ado...

Aug 21 2026 2608.20639v1
Privacy-Preserving Object Detection for Vision Transformer-Based Models

We propose a novel object detection method that enables us to protect sensitive visual information of test images. Previous studies considering visual...

Aug 21 2026 2608.20712v1
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs

Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, ...

Aug 21 2026 2608.20999v1
Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset

Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs, which use...

Aug 21 2026 2608.21189v1
Parametric neural control differentiates top neural network models of primate visual cortex

Leading deep neural network encoding models predict visual cortical responses with nearly indistinguishable accuracy, raising the strong inference tha...

Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment

Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI s...

Aug 20 2026 2608.19825v1
Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal

High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured b...

Aug 20 2026 2608.20212v1
GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual e...

Aug 19 2026 2608.19355v1
General-purpose time-series foundation models enable sample-efficient transfer learning in retinal electrophysiology

Electroretinography (ERG) measures the functional response of distinct retinal cells to light, but was largely displaced by structural imaging in the ...

OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation

Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistently high per...

Aug 19 2026 2608.18516v1
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection

Diffusion-based generators have made synthetic images ubiquitous, but detectors often fail under simultaneous shifts in generator, prompt/style, and s...

Aug 19 2026 2608.18523v1
When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models

Aligned vision-language models (VLMs) are designed to balance grounded visual reasoning with safe generation behavior. However, we observe a striking ...

Aug 19 2026 2608.18628v1
Browse Categories