Latest AI and machine learning research in ophthalmology for healthcare professionals.
Multimodal Large Language Models (MLLMs) have achieved notable performance in computer vision tasks that require reasoning across visual and textual modalities, yet their capabilities are limited to their pre-trained data, requiring extensive fine-tuning for updates. Recent researches have explored the use of In-Context Learning (ICL) to overcome these challenges by providing a set of demonstrat...
We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance computing (HPC) systems. MOFA addresses key challenges in integrating GPU-accelerated computing for GPU-intensive GenAI tasks, including distributed training and inference, alongside CPU- and GPU-optimized tasks for scree...
Convolutional Neural Networks (CNN) and Vision Transformers (ViT) have dominated the field of Computer Vision (CV). Graph Neural Networks (GNN) have...
Automotive Simulation is a potentially cost-effective strategy to identify and test corner case scenarios in automotive perception. Recent work has ...
This paper investigates the robustness of vision-language models against adversarial visual perturbations and introduces a novel ``double visual def...
Few-shot learning in medical image classification presents a significant challenge due to the limited availability of annotated data and the complex...
Kolmogorov-Arnold Neural Networks (KANs) have gained significant attention in the machine learning community. However, their implementation often su...
Vision-based tactile sensors have drawn increasing interest in the robotics community. However, traditional lens-based designs impose minimum thickn...
The HyperText Markup Language 5 (HTML5)
Single image defocus deblurring (SIDD) aims to restore an all-in-focus image from a defocused one. Distribution shifts in defocused images generally...
Implementing accurate models of the retina is a challenging task, particularly in the context of creating visual prosthetics and devices. Notwithsta...
Remote sensing imagery is dense with objects and contextual visual information. There is a recent trend to combine paired satellite images and text ...
In the current era of Machine Learning, Transformers have become the de facto approach across a variety of domains, such as computer vision and natu...
Large Language Models (LLMs) have shown impressive potential in clinical question answering (QA), with Retrieval Augmented Generation (RAG) emerging...
Birds Eye View perception models require extensive data to perform and generalize effectively. While traditional datasets often provide abundant dri...
As trends in education evolve, personalized learning has transformed individuals' engagement with knowledge and skill development. In the digital ag...
Sclera segmentation is crucial for developing automatic eye-related medical computer-aided diagnostic systems, as well as for personal identificatio...
Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a sign...
The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biome...
Compositional Zero-Shot Learning (CZSL) aims to enable models to recognize novel compositions of visual states and objects that were absent during t...