Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4861-4880 of 9,853 articles

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Multimodal foundation models, such as GPT-4o, have recently made remarkable progress, but it is not clear where exactly these models stand in terms of understanding vision. In this paper, we benchmark the performance of popular multimodal foundation models (GPT-4o, o4-mini, Gemini 1.5 Pro and Gemini 2.0 Flash, Claude 3.5 Sonnet, Qwen2-VL, Llama 3.2) on standard computer vision tasks (semantic se...

Characterizing control between interacting subsystems with deep Jacobian estimation

Biological function arises through the dynamical interactions of multiple subsystems, including those between brain areas, within gene regulatory networks, and more. A common approach to understanding these systems is to model the dynamics of each subsystem and characterize communication between them. An alternative approach is through the lens of control theory: how the subsystems control one a...

evMLP: An Efficient Event-Driven MLP Architecture for Vision

Deep neural networks have achieved remarkable results in computer vision tasks. In the early days, Convolutional Neural Networks (CNNs) were the mai...

When Does Pruning Benefit Vision Representations?

Pruning is widely used to reduce the complexity of deep learning models, but its effects on interpretability and representation learning remain poor...

Vision-Aided ISAC in Low-Altitude Economy Networks via De-Diffused Visual Priors

Emerging low-altitude economy networks (LAENets) require agile and privacy-preserving resource control under dynamic agent mobility and limited infr...

Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation

We introduce Generalized Test-Time Augmentation (GTTA), a highly effective method for improving the performance of a trained model, which unlike oth...

Evaluating Large Language Models for Multimodal Simulated Ophthalmic Decision-Making in Diabetic Retinopathy and Glaucoma Screening

Large language models (LLMs) can simulate clinical reasoning based on natural language prompts, but their utility in ophthalmology is largely unexpl...

AI Meets Maritime Training: Precision Analytics for Enhanced Safety and Performance

Traditional simulator-based training for maritime professionals is critical for ensuring safety at sea but often depends on subjective trainer asses...

Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks durin...

Research on Improving the High Precision and Lightweight Diabetic Retinopathy Detection of YOLOv8n

Early detection and diagnosis of diabetic retinopathy is one of the current research focuses in ophthalmology. However, due to the subtle features o...

Stable Tracking of Eye Gaze Direction During Ophthalmic Surgery

Ophthalmic surgical robots offer superior stability and precision by reducing the natural hand tremors of human surgeons, enabling delicate operatio...

BadViM: Backdoor Attack against Vision Mamba

Vision State Space Models (SSMs), particularly architectures like Vision Mamba (ViM), have emerged as promising alternatives to Vision Transformers ...

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

The architecture of multimodal large language models (MLLMs) commonly connects a vision encoder, often based on CLIP-ViT, to a large language model....

Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models

Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primaril...

Just Noticeable Difference for Large Multimodal Models

Just noticeable difference (JND), the minimum change that the human visual system (HVS) can perceive, has been studied for decades. Although recent ...

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and actin...

Fundus Refraction Offset as a Personalized Biomarker for 12-Year Risk of Retinal Detachment.

PURPOSE: The purpose of this study was to investigate the potential of a novel anatomical metric of ametropia-fundus refraction offset (FRO)-in strati...

Jul 1 2025 40590806
Coal classification and analysis based on shadowgraphy and deep learning methods.

The classification and analysis of coal are crucial for energy production and resource management. Shadowgraphy, leveraging variations in air refracti...

Jul 1 2025 40591303
Slice-Inference-Assisted Lightweight Small Object Detection Model for Holographic Digital Immunoassay Quantification.

Sensitive and cost-effective detection methods utilizing portable equipment are crucial for applications in food safety inspection, environmental moni...

Jul 1 2025 40540441
Automatic transformer-based grading of multiple retinal inflammatory signs in uveitis on fluorescein angiography.

BACKGROUND: Grading fluorescein angiography (FA) for uveitis is complex, often leading to the oversight of retinal inflammation in clinical studies. T...

Jul 1 2025 40403640
Browse Categories