Latest AI and machine learning research in ophthalmology for healthcare professionals.
Can computer vision help us explore the ocean? The ultimate challenge for computer vision is to recognize any visual phenomena, more than only the objects and animals humans encounter in their terrestrial lives. Previous datasets have explored everyday objects and fine-grained categories humans see frequently. We present the FathomVerse v0 detection dataset to push the limits of our field by exp...
Vision-based tactile sensors, through high-resolution optical measurements, can effectively perceive the geometric shape of objects and the force information during the contact process, thus helping robots acquire higher-dimensional tactile data. Vision-based tactile sensor simulation supports the acquisition and understanding of tactile information without physical sensors by accurately capturi...
Spatial understanding of the semantics of the surroundings is a key capability needed by autonomous cars to enable safe driving decisions. Recently,...
Learning lighting adaption is a key step in obtaining a good visual perception and supporting downstream vision tasks. There are multiple light-rela...
Quality degradation is observed in underwater images due to the effects of light refraction and absorption by water, leading to issues like color ca...
Museums serve as vital repositories of cultural heritage and historical artifacts spanning diverse epochs, civilizations, and regions, preserving we...
Visual framing analysis is a key method in social sciences for determining common themes and concepts in a given discourse. To reduce manual effort,...
Multimodal LLMs (MLLMs) equip language models with visual capabilities by aligning vision encoders with language models. Existing methods to enhance...
Most existing GUI agents typically depend on non-vision inputs like HTML source code or accessibility trees, limiting their flexibility across diver...
The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applicatio...
PURPOSE: Eye rubbing is considered to play a significant role in the progression of keratoconus and of corneal ectasia following refractive surgery. T...
PURPOSE: Identify optimal metabolic features and pathways across diabetic retinopathy (DR) stages, develop risk models to differentiate diabetic macul...
Large Vision Language Models (LVLMs) have achieved remarkable performance in various vision-language tasks. However, it is still unclear how accurat...
Industrial anomaly detection (IAD) plays a crucial role in the maintenance and quality control of manufacturing processes. In this paper, we propose...
The zero-shot performance of object detectors degrades when tested on different modalities, such as infrared and depth. While recent work has explor...
PURPOSE: To differentiate a normal cornea from a forme fruste keratoconus (FFKC) with the swept-source optical coherence tomography (SS-OCT) topograph...
The retinogeniculate visual pathway (RGVP) is responsible for carrying visual information from the retina to the lateral geniculate nucleus. Identific...
BACKGROUND: Age-related macular degeneration (AMD) is a major cause of irreversible visual impairment, with dry AMD being the most prevalent form. Pro...
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-worl...
The increasing proportion of the older adult population has made the smart home care industry one of the critical markets for virtual human-like age...