Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4821-4840 of 9,853 articles

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models

Prompt learning facilitates the efficient adaptation of Vision-Language Models (VLMs) to various downstream tasks. However, it faces two significant challenges: (1) inadequate modeling of class embedding distributions for unseen instances, leading to suboptimal generalization on novel classes; (2) prevailing methodologies predominantly confine cross-modal alignment to the final output layer of v...

From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning

Low-level enhancement and high-level visual understanding in low-light vision have traditionally been treated separately. Low-light enhancement improves image quality for downstream tasks, but existing methods rely on physical or geometric priors, limiting generalization. Evaluation mainly focuses on visual quality rather than downstream performance. Low-light visual understanding, constrained b...

Formation and Regulation of Calcium Sparks on a Nonlinear Spatial Network of Ryanodine Receptors

Accurate regulation of calcium release is essential for cellular signaling, with the spatial distribution of ryanodine receptors (RyRs) playing a cr...

[Problems and countermeasures in eye care and vision screening services for children aged 0 to 6 years].

The critical period for visual function and ocular structure development occurs from 0 to 6 years of age, making standardized eye care and vision scre...

Jul 11 2025 40605299
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, ...

Hardware-Aware Feature Extraction Quantisation for Real-Time Visual Odometry on FPGA Platforms

Accurate position estimation is essential for modern navigation systems deployed in autonomous platforms, including ground vehicles, marine vessels,...

Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought

Large vision-language models (LVLMs) have demonstrated remarkable capabilities by integrating pre-trained vision encoders with large language models...

ViLU: Learning Vision-Language Uncertainties for Failure Prediction

Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new V...

ViLU: Learning Vision-Language Uncertainties for Failure Prediction

Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new V...

Synthetic MC via Biological Transmitters: Therapeutic Modulation of the Gut-Brain Axis

Synthetic molecular communication (SMC) is a key enabler for future healthcare systems in which Internet of Bio-Nano-Things (IoBNT) devices facilita...

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

Building state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-qual...

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

Building state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-qual...

Robust Containerization of the High Angular Resolution Functional Imaging (HARFI) Pipeline

Historically, functional magnetic resonance imaging (fMRI) of the brain has focused primarily on gray matter, particularly the cortical gray matter ...

VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation

Graphical User Interface (GUI) agents powered by Large Vision-Language Models (LVLMs) have emerged as a revolutionary approach to automating human-m...

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

Technical reports and articles often contain valuable information in the form of semi-structured data like charts, and figures. Interpreting these a...

Omni-Video: Democratizing Unified Video Understanding and Generation

Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, produc...

Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023

Search engines like Google have become major sources of information for voters during election campaigns. To assess potential biases across candidat...

Event-RGB Fusion for Spacecraft Pose Estimation Under Harsh Lighting

Spacecraft pose estimation is crucial for autonomous in-space operations, such as rendezvous, docking and on-orbit servicing. Vision-based pose esti...

R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding

Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by ad...

Simulating Refractive Distortions and Weather-Induced Artifacts for Resource-Constrained Autonomous Perception

The scarcity of autonomous vehicle datasets from developing regions, particularly across Africa's diverse urban, rural, and unpaved roads, remains a...

Browse Categories