Latest AI and machine learning research in ophthalmology for healthcare professionals.
The Vision Transformer (ViT) has made significant advancements in computer vision, utilizing self-attention mechanisms to achieve state-of-the-art performance across various tasks, including image classification, object detection, and segmentation. Its architectural flexibility and capabilities have made it a preferred choice among researchers and practitioners. However, the intricate multi-head...
Diabetic retinopathy (DR), a serious ocular complication of diabetes, is one of the primary causes of vision loss among retinal vascular diseases. Deep learning methods have been extensively applied in the grading of diabetic retinopathy (DR). However, their performance declines significantly when applied to data outside the training distribution due to domain shifts. Domain generalization (DG) ...
Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from seq...
Single-Domain Generalized Object Detection~(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse un...
Spontaneous waves are ubiquitous during early brain development and are hypothesized to drive the development of receptive fields (RFs). Different s...
Data augmentation is essential in medical imaging for improving classification accuracy, lesion detection, and organ segmentation under limited data...
Purpose: Automated Surgical Phase Recognition (SPR) uses Artificial Intelligence (AI) to segment the surgical workflow into its key events, function...
Compression has been a critical lens to understand the success of Transformers. In the past, we have typically taken the target distribution as a cr...
Mild Traumatic Brain Injury (TBI) detection presents significant challenges due to the subtle and often ambiguous presentation of symptoms in medica...
Deep neural networks (DNNs) have proven to be successful in various computer vision applications such that models even infer in safety-critical situ...
With the surge of large language models (LLMs), Large Vision-Language Models (VLMs)--which integrate vision encoders with LLMs for accurate visual g...
Molecular Communication (MC) has long been envisioned to enable an Internet of Bio-Nano Things (IoBNT) with medical applications, where nanomachines...
In an era where social media platforms abound, individuals frequently share images that offer insights into their intents and interests, impacting i...
Large-scale models trained on extensive datasets have become the standard due to their strong generalizability across diverse tasks. In-context lear...
In the field of medical imaging, the advent of deep learning, especially the application of convolutional neural networks (CNNs) has revolutionized ...
Ophthalmic diseases pose a significant global health challenge, yet traditional diagnosis methods and existing single-eye deep learning approaches o...
Vision Transformers (ViTs) excel in semantic segmentation but demand significant computation, posing challenges for deployment on resource-constrain...
Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information...
In the field of image recognition, spiking neural networks (SNNs) have achieved performance comparable to conventional artificial neural networks (A...
In the field of image recognition, spiking neural networks (SNNs) have achieved performance comparable to conventional artificial neural networks (A...