Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5221-5240 of 9,853 articles

Visual Agentic Reinforcement Fine-Tuning

A key trend in Large Reasoning Models (e.g., OpenAI's o3) is the native agentic ability to use external tools such as web browsers for searching and writing/executing code for image manipulation to think with images. In the open-source research community, while significant progress has been made in language-only agentic abilities such as function calling and tool integration, the development of ...

VoQA: Visual-only Question Answering

We propose Visual-only Question Answering (VoQA), a novel multimodal task in which questions are visually embedded within images, without any accompanying textual input. This requires models to locate, recognize, and reason over visually embedded textual questions, posing challenges for existing large vision-language models (LVLMs), which show notable performance drops even with carefully design...

Scaling Vision Mamba Across Resolutions via Fractal Traversal

Vision Mamba has recently emerged as a promising alternative to Transformer-based architectures, offering linear complexity in sequence length while...

[Expert consensus on systematic assessment and treatment of refractory wounds in the elderly (2025 edition)].

Early prevention and standardized management of refractory wounds in the elderly are very important for improving prognosis, reducing disability rate,...

May 20 2025 40419353
Real-time driver drowsiness detection using transformer architectures: a novel deep learning approach.

Driver drowsiness is a leading cause of road accidents, resulting in significant societal, economic, and emotional losses. This paper introduces a nov...

May 20 2025 40394076
The Application of Artificial Intelligence to Enhance Spinal Cord Stimulation Efficacy for Chronic Pain Management: Current Evidence and Future Directions.

PURPOSE OF REVIEW: Chronic pain significantly impacts quality of life for millions globally, with spinal cord stimulation (SCS) as an established trea...

May 20 2025 40394275
I'll believe it when I see it: Images increase misinformation sharing in Vision-Language Models

Large language models are increasingly integrated into news recommendation systems, raising concerns about their role in spreading misinformation. I...

Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts

With the rapid improvement of machine learning (ML) models, cognitive scientists are increasingly asking about their alignment with how humans think...

Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision

Inspired by foveal vision, hard attention models promise interpretability and parameter economy. However, existing models like the Recurrent Model o...

Combinatorial Sample-and Back-Focal-Plane (BFP) Imaging. Pt. I: Instrument and acquisition parameters affecting BFP images and their analysis

The back-focal plane (BFP) of a high-numerical aperture objective contains the fluoro-phore radiation pattern, which encodes information about the a...

ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models

Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans compared to planning i...

3D Visual Illusion Depth Estimation

3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, maki...

RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions

The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models fo...

Mamba-Adaptor: State Space Model Adaptor for Visual Recognition

Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency...

Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps

Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-com...

MVPainter: Accurate and Detailed 3D Texture Generation via Multi-View Diffusion with Geometric Control

Recently, significant advances have been made in 3D object generation. Building upon the generated geometry, current pipelines typically employ imag...

Fast and accurate active alignment of camera lenses with physics-informed deep learning.

As optical systems become increasingly complex, accurate and fast alignment is becoming more critical. Active alignment (AA) techniques dynamically op...

May 19 2025 40515030
A CNN-based approach for detecting eye blink episodes in EEG signals.

This study aims to develop and evaluate a convolutional neural network (CNN)-based architecture for detecting eye blink episodes in electroencephalogr...

May 19 2025 40328273
Kornia-rs: A Low-Level 3D Computer Vision Library In Rust

We present \textit{kornia-rs}, a high-performance 3D computer vision library written entirely in native Rust, designed for safety-critical and real-...

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding

Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of s...

Browse Categories