Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5621-5640 of 9,853 articles

Panorama Generation From NFoV Image Done Right

Generating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is \textbf{not suitable for evaluating the distortion}. In this work, we first propose a distortion-specific CL...

VTD-CLIP: Video-to-Text Discretization via Prompting CLIP

Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-efficient fine-tuning of image-text pre-trained models, yet they often suffer from limited interpretability and poor generalization due to inadequate temporal modeling. To address these, we propose a simple yet effective vid...

Efficient Deep Learning Approaches for Processing Ultra-Widefield Retinal Imaging

Deep learning has emerged as the predominant solution for classifying medical images. We intend to apply these developments to the ultra-widefield (...

Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook

Retrieval-augmented generation (RAG) has emerged as a pivotal technique in artificial intelligence (AI), particularly in enhancing the capabilities ...

Guided Diffusion for the Extension of Machine Vision to Human Visual Perception

Image compression technology eliminates redundant information to enable efficient transmission and storage of images, serving both machine vision an...

FundusGAN: A Hierarchical Feature-Aware Generative Framework for High-Fidelity Fundus Image Generation

Recent advancements in ophthalmology foundation models such as RetFound have demonstrated remarkable diagnostic capabilities but require massive dat...

Visual Variational Autoencoder Prompt Tuning

Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large vision transformers to downstream tasks without the proh...

AI-Based Screening for Depression and Social Anxiety Through Eye Tracking: An Exploratory Study

Well-being is a dynamic construct that evolves over time and fluctuates within individuals, presenting challenges for accurate quantification. Reduc...

OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement

Recent advancements demonstrated by DeepSeek-R1 have shown that complex reasoning abilities in large language models (LLMs), including sophisticated...

A Deep Learning Framework for Visual Attention Prediction and Analysis of News Interfaces

News outlets' competition for attention in news interfaces has highlighted the need for demographically-aware saliency prediction models. Despite re...

Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

Vision-Language Models (VLMs) learn a shared feature space for text and images, enabling the comparison of inputs of different modalities. While pri...

Beyond Accuracy: What Matters in Designing Well-Behaved Models?

Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they o...

RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark

In the computer vision community, the preference for pre-training visual models has largely shifted toward sRGB images due to their ease of acquisit...

EasyRobust: A Comprehensive and Easy-to-use Toolkit for Robust and Generalized Vision

Deep neural networks (DNNs) has shown great promise in computer vision tasks. However, machine vision achieved by DNNs cannot be as robust as human ...

Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning

Vision Language Models exhibited immense potential for embodied AI, yet they often lack the sophisticated situational reasoning required for complex...

When Less is Enough: Adaptive Token Reduction for Efficient Image Representation

Vision encoders typically generate a large number of visual tokens, providing information-rich representations but significantly increasing computat...

Do Visual Imaginations Improve Vision-and-Language Navigation Agents?

Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we s...

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In part...

Bokehlicious: Photorealistic Bokeh Rendering with Controllable Apertures

Bokeh rendering methods play a key role in creating the visually appealing, softly blurred backgrounds seen in professional photography. While recen...

Beyond the Visible: Multispectral Vision-Language Learning for Earth Observation

Vision-language models for Earth observation (EO) typically rely on the visual spectrum of data as the only model input, thus failing to leverage th...

Browse Categories