Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6021-6040 of 9,853 articles

Attribute-based Visual Reprogramming for Vision-Language Models

Visual reprogramming (VR) reuses pre-trained vision models for downstream image classification tasks by adding trainable noise patterns to inputs. When applied to vision-language models (e.g., CLIP), existing VR approaches follow the same pipeline used in vision models (e.g., ResNet, ViT), where ground-truth class labels are inserted into fixed text templates to guide the optimization of VR patt...

Learning Visual Proxy for Compositional Zero-Shot Learning

Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Existing methods align textual prototypes with visual features through Vision-Language Models (VLMs), but they face two key limitations: (1) modality gaps hinder the discrimination of semantically similar composition pairs, and (2) single-modal textual pro...

Enhancing Medical Image Analysis through Geometric and Photometric transformations

Medical image analysis suffers from a lack of labeled data due to several challenges including patient privacy and lack of experts. Although some AI...

A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs

A fundamental challenge in artificial intelligence involves understanding the cognitive mechanisms underlying visual reasoning in sophisticated mode...

Patch-Based and Non-Patch-Based inputs Comparison into Deep Neural Models: Application for the Segmentation of Retinal Diseases on Optical Coherence Tomography Volumes

Worldwide, sight loss is commonly occurred by retinal diseases, with age-related macular degeneration (AMD) being a notable facet that affects elder...

UniRestore: Unified Perceptual and Task-Oriented Image Restoration Model Using Diffusion Prior

Image restoration aims to recover content from inputs degraded by various factors, such as adverse weather, blur, and noise. Perceptual Image Restor...

Need for Speed: A Comprehensive Benchmark of JPEG Decoders in Python

Image loading represents a critical bottleneck in modern machine learning pipelines, particularly in computer vision tasks where JPEG remains the do...

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

In this paper, we propose VideoLLaMA3, a more advanced multimodal foundation model for image and video understanding. The core design philosophy of ...

SMART-Vision: Survey of Modern Action Recognition Techniques in Vision

Human Action Recognition (HAR) is a challenging domain in computer vision, involving recognizing complex patterns by analyzing the spatiotemporal dy...

PreciseCam: Precise Camera Control for Text-to-Image Generation

Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is...

Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models

Animal stereotypes are deeply embedded in human culture and language. They often shape our perceptions and expectations of various species. Our stud...

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

We present VARGPT, a novel multimodal large language model (MLLM) that unifies visual understanding and generation within a single autoregressive fr...

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities in understanding and describing visual content, achieving state-of-th...

Adaptive Class Learning to Screen Diabetic Disorders in Fundus Images of Eye

The prevalence of ocular illnesses is growing globally, presenting a substantial public health challenge. Early detection and timely intervention ar...

Are Traditional Deep Learning Model Approaches as Effective as a Retinal-Specific Foundation Model for Ocular and Systemic Disease Detection?

Background: RETFound, a self-supervised, retina-specific foundation model (FM), showed potential in downstream applications. However, its comparativ...

WaveNet-SF: A Hybrid Network for Retinal Disease Detection Based on Wavelet Transform in the Spatial-Frequency Domain

Retinal diseases are a leading cause of vision impairment and blindness, with timely diagnosis being critical for effective treatment. Optical Coher...

Transferability of labels between multilens cameras

In this work, a new method for automatically extending Bounding Box (BB) and mask labels across different channels on multilens cameras is presented...

Anomaly Detection for Industrial Applications, Its Challenges, Solutions, and Future Directions: A Review

Anomaly detection from images captured using camera sensors is one of the mainstream applications at the industrial level. Particularly, it maintain...

DeepEyeNet: Adaptive Genetic Bayesian Algorithm Based Hybrid ConvNeXtTiny Framework For Multi-Feature Glaucoma Eye Diagnosis

Glaucoma is a leading cause of irreversible blindness worldwide, emphasizing the critical need for early detection and intervention. In this paper, ...

Generative Physical AI in Vision: A Survey

Generative Artificial Intelligence (AI) has rapidly advanced the field of computer vision by enabling machines to create and interpret visual data w...

Browse Categories