Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6041-6060 of 9,853 articles

Visual RAG: Expanding MLLM visual knowledge without fine-tuning

Multimodal Large Language Models (MLLMs) have achieved notable performance in computer vision tasks that require reasoning across visual and textual modalities, yet their capabilities are limited to their pre-trained data, requiring extensive fine-tuning for updates. Recent researches have explored the use of In-Context Learning (ICL) to overcome these challenges by providing a set of demonstrat...

MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow

We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance computing (HPC) systems. MOFA addresses key challenges in integrating GPU-accelerated computing for GPU-intensive GenAI tasks, including distributed training and inference, alongside CPU- and GPU-optimized tasks for scree...

ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning

Convolutional Neural Networks (CNN) and Vision Transformers (ViT) have dominated the field of Computer Vision (CV). Graph Neural Networks (GNN) have...

SOLAS: Superpositioning an Optical Lens in Automotive Simulation

Automotive Simulation is a potentially cost-effective strategy to identify and test corner case scenarios in automotive perception. Recent work has ...

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness

This paper investigates the robustness of vision-language models against adversarial visual perturbations and introduces a novel ``double visual def...

Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning

Few-shot learning in medical image classification presents a significant challenge due to the limited availability of annotated data and the complex...

Free-Knots Kolmogorov-Arnold Network: On the Analysis of Spline Knots and Advancing Stability

Kolmogorov-Arnold Neural Networks (KANs) have gained significant attention in the machine learning community. However, their implementation often su...

ThinTact:Thin Vision-Based Tactile Sensor by Lensless Imaging

Vision-based tactile sensors have drawn increasing interest in the robotics community. However, traditional lens-based designs impose minimum thickn...

Exploring the Capabilities of Vision-Language Models to Detect Visual Bugs in HTML5 Applications

The HyperText Markup Language 5 (HTML5) is useful for creating visual-centric web applications. However, unlike traditional web application...

Continual Test-Time Adaptation for Single Image Defocus Deblurring via Causal Siamese Networks

Single image defocus deblurring (SIDD) aims to restore an all-in-focus image from a defocused one. Distribution shifts in defocused images generally...

Neuromorphic Retina: An FPGA-based Emulator

Implementing accurate models of the retina is a challenging task, particularly in the context of creating visual prosthetics and devices. Notwithsta...

FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing

Remote sensing imagery is dense with objects and contextual visual information. There is a recent trend to combine paired satellite images and text ...

Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time

In the current era of Machine Learning, Transformers have become the de facto approach across a variety of domains, such as computer vision and natu...

ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems

Large Language Models (LLMs) have shown impressive potential in clinical question answering (QA), with Retrieval Augmented Generation (RAG) emerging...

Revisiting Birds Eye View Perception Models with Frozen Foundation Models: DINOv2 and Metric3Dv2

Birds Eye View perception models require extensive data to perform and generalize effectively. While traditional datasets often provide abundant dri...

Assessment of Personalized Learning in Immersive and Intelligent Virtual Classroom on Student Engagement

As trends in education evolve, personalized learning has transformed individuals' engagement with knowledge and skill development. In the digital ag...

Boosting Sclera Segmentation through Semi-supervised Learning with Fewer Labels

Sclera segmentation is crucial for developing automatic eye-related medical computer-aided diagnostic systems, as well as for personal identificatio...

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a sign...

BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biome...

Duplex: Dual Prototype Learning for Compositional Zero-Shot Learning

Compositional Zero-Shot Learning (CZSL) aims to enable models to recognize novel compositions of visual states and objects that were absent during t...

Browse Categories