Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5061-5080 of 9,853 articles

A Compendium of Autonomous Navigation using Object Detection and Tracking in Unmanned Aerial Vehicles

Unmanned Aerial Vehicles (UAVs) are one of the most revolutionary inventions of 21st century. At the core of a UAV lies the central processing system that uses wireless signals to control their movement. The most popular UAVs are quadcopters that use a set of four motors, arranged as two on either side with opposite spin. An autonomous UAV is called a drone. Drones have been in service in the US...

HueManity: Probing Fine-Grained Visual Perception in MLLMs

Multimodal Large Language Models (MLLMs) excel at high-level visual reasoning, but their performance on nuanced perceptual tasks remains surprisingly limited. We present HueManity, a benchmark designed to assess visual perception in MLLMs. The dataset comprises 83,850 images featuring two-character alphanumeric strings embedded in Ishihara test style dot patterns, challenging models on precise p...

Browser Fingerprinting Using WebAssembly

Web client fingerprinting has become a widely used technique for uniquely identifying users, browsers, operating systems, and devices with high accu...

Fovea Stacking: Imaging with Dynamic Localized Aberration Correction

The desire for cameras with smaller form factors has recently lead to a push for exploring computational imaging systems with reduced optical comple...

Text-to-CT Generation via 3D Latent Diffusion Model with Contrastive Vision-Language Pretraining

Objective: While recent advances in text-conditioned generative models have enabled the synthesis of realistic medical images, progress has been lar...

Flying Co-Stereo: Enabling Long-Range Aerial Dense Mapping via Collaborative Stereo Vision of Dynamic-Baseline

Lightweight long-range mapping is critical for safe navigation of UAV swarms in large-scale unknown environments. Traditional stereo vision systems ...

Imputation of Missing Data in Smooth Pursuit Eye Movements Using a Self-Attention-based Deep Learning Approach

Missing data is a relevant issue in time series, especially in biomedical sequences such as those corresponding to smooth pursuit eye movements, whi...

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successf...

Development and validation of a 3-D deep learning system for diabetic macular oedema classification on optical coherence tomography images.

OBJECTIVES: To develop and validate an automated diabetic macular oedema (DME) classification system based on the images from different three-dimensio...

May 31 2025 40449950
Multi-Analyte, Swab-based Automated Wound Monitor with AI

Diabetic foot ulcers (DFUs), a class of chronic wounds, affect ~750,000 individuals every year in the US alone and identifying non-healing DFUs that...

Lightweight Convolutional Neural Networks for Retinal Disease Classification

Retinal diseases such as Diabetic Retinopathy (DR) and Macular Hole (MH) significantly impact vision and affect millions worldwide. Early detection ...

PerFormer: A Permutation Based Vision Transformer for Remaining Useful Life Prediction

Accurately estimating the remaining useful life (RUL) for degradation systems is crucial in modern prognostic and health management (PHM). Convoluti...

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. Ho...

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL

Although chain-of-thought reasoning and reinforcement learning (RL) have driven breakthroughs in NLP, their integration into generative vision model...

Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck

This paper reveals that many state-of-the-art large language models (LLMs) lack hierarchical knowledge about our visual world, unaware of even well-...

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts

Multimodal large language models (MLLMs) require a nuanced interpretation of complex image information, typically leveraging a vision encoder to per...

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, r...

KEVER^2: Knowledge-Enhanced Visual Emotion Reasoning and Retrieval

Understanding what emotions images evoke in their viewers is a foundational goal in human-centric visual computing. While recent advances in vision-...

Harnessing Foundation Models for Robust and Generalizable 6-DOF Bronchoscopy Localization

Vision-based 6-DOF bronchoscopy localization offers a promising solution for accurate and cost-effective interventional guidance. However, existing ...

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models

While adversarial attacks on vision-and-language pretraining (VLP) models have been explored, generating natural adversarial samples crafted through...

Browse Categories