Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6321-6340 of 9,853 articles

HyperCLIP: Adapting Vision-Language models with Hypernetworks

Self-supervised vision-language models trained with contrastive objectives form the basis of current state-of-the-art methods in AI vision tasks. The success of these models is a direct consequence of the huge web-scale datasets used to train them, but they require correspondingly large vision components to properly learn powerful and general representations from such a broad data domain. This p...

UNEM: UNrolled Generalized EM for Transductive Few-Shot Learning

Transductive few-shot learning has recently triggered wide attention in computer vision. Yet, current methods introduce key hyper-parameters, which control the prediction statistics of the test batches, such as the level of class balance, affecting performances significantly. Such hyper-parameters are empirically grid-searched over validation data, and their configurations may vary substantially...

Parameterized Complexity of Caching in Networks

The fundamental caching problem in networks asks to find an allocation of contents to a network of caches with the aim of maximizing the cache hit r...

Sensitive Image Classification by Vision Transformers

When it comes to classifying child sexual abuse images, managing similar inter-class correlations and diverse intra-class correlations poses a signi...

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding

Flowcharts are typically presented as images, driving the trend of using vision-language models (VLMs) for end-to-end flowchart understanding. Howev...

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. Howeve...

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

The rapid advance of Large Language Models (LLMs) has catalyzed the development of Vision-Language Models (VLMs). Monolithic VLMs, which avoid modal...

Demystifying the Potential of ChatGPT-4 Vision for Construction Progress Monitoring

The integration of Large Vision-Language Models (LVLMs) such as OpenAI's GPT-4 Vision into various sectors has marked a significant evolution in the...

CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images

3D Gaussian Splatting (3DGS) has attracted significant attention for its high-quality novel view rendering, inspiring research to address real-world...

Mamba-based Deep Learning Approaches for Sleep Staging on a Wireless Multimodal Wearable System without Electroencephalography

Study Objectives: We investigate Mamba-based deep learning approaches for sleep staging on signals from ANNE One (Sibel Health, Evanston, IL), a non...

NeuroPump: Simultaneous Geometric and Color Rectification for Underwater Images

Underwater image restoration aims to remove geometric and color distortions due to water refraction, absorption and scattering. Previous studies foc...

Computing the Non-Dominated Flexible Skyline in Vertically Distributed Datasets with No Random Access

In today's data-driven world, algorithms operating with vertically distributed datasets are crucial due to the increasing prevalence of large-scale,...

LG-Sleep: Local and Global Temporal Dependencies for Mice Sleep Scoring

Efficiently identifying sleep stages is crucial for unraveling the intricacies of sleep in both preclinical and clinical research. The labor-intensi...

PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation

Despite significant advancements in Large Vision-Language Models (LVLMs), existing pixel-grounding models operate on single-image settings, limiting...

Leveraging Color Channel Independence for Improved Unsupervised Object Detection

Object-centric architectures can learn to extract distinct object representations from visual scenes, enabling downstream applications on the object...

TDCNet: Transparent Objects Depth Completion with CNN-Transformer Dual-Branch Parallel Network

The sensing and manipulation of transparent objects present a critical challenge in industrial and laboratory robotics. Conventional sensors face ch...

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

Large Vision Language Models (LVLMs) have achieved significant progress in integrating visual and textual inputs for multimodal reasoning. However, ...

Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning

Few-shot, fine-grained classification in computer vision poses significant challenges due to the need to differentiate subtle class distinctions wit...

Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers

Due to its efficiency, Post-Training Quantization (PTQ) has been widely adopted for compressing Vision Transformers (ViTs). However, when quantized ...

LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining

Visual Information Extraction (VIE) plays a crucial role in the comprehension of semi-structured documents, and several pre-trained models have been...

Browse Categories