Latest AI and machine learning research in ophthalmology for healthcare professionals.
Implicit neural representations (INRs), which leverage neural networks to represent signals by mapping coordinates to their corresponding attributes, have garnered significant attention. They are extensively utilized for image representation, with pixel coordinates as input and pixel values as output. In contrast to prior works focusing on investigating the effect of the model's inside component...
While deep learning has exhibited remarkable predictive capabilities in various medical image tasks, its inherent black-box nature has hindered its widespread implementation in real-world healthcare settings. Our objective is to unveil the decision-making processes of deep learning models in the context of glaucoma classification by employing several Class Activation Map (CAM) techniques to gene...
Modern AI models excel in controlled settings but often fail in real-world scenarios where data distributions shift unpredictably - a challenge know...
Recently, vision transformers were shown to be capable of outperforming convolutional neural networks when pretrained on sufficiently large datasets...
Subretinal injection is a critical procedure for delivering therapeutic agents to treat retinal diseases such as age-related macular degeneration (A...
Multi-label Recognition (MLR) involves assigning multiple labels to each data instance in an image, offering advantages over single-label classifica...
Diabetic Retinopathy (DR) is a serious and common complication of diabetes, caused by prolonged high blood sugar levels that damage the small retina...
Grasping is a fundamental skill for interacting with the environment. However, this ability can be difficult for some (e.g. due to disability). Wear...
Vision Foundation Models (VFMs) and Vision-Language Models (VLMs) have gained traction in Domain Generalized Semantic Segmentation (DGSS) due to the...
Conventional Vision-Language Models(VLMs) typically utilize a fixed number of vision tokens, regardless of task complexity. This one-size-fits-all s...
As AI takes on increasingly complex roles in human-computer interaction, fundamental questions arise: how can HCI help maintain the user as the prim...
In Visual Document Understanding (VDU) tasks, fine-tuning a pre-trained Vision-Language Model (VLM) with new datasets often falls short in optimizin...
Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However...
Transformer-based architectures have revolutionized the landscape of deep learning. In computer vision domain, Vision Transformer demonstrates remar...
This paper presents the Sensorimotor Transformer (SMT), a vision model inspired by human saccadic eye movements that prioritize high-saliency region...
Efficiently understanding long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms fo...
Many images and videos are primarily processed by computer vision algorithms, involving only occasional human inspection. When this content requires...
Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significan...
Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES...
Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are lim...