Latest AI and machine learning research in ophthalmology for healthcare professionals.
This study investigates the performance of the two most relevant computer vision deep learning architectures, Convolutional Neural Network and Vision Transformer, for event-based cameras. These cameras capture scene changes, unlike traditional frame-based cameras with capture static images, and are particularly suited for dynamic environments such as UAVs and autonomous vehicles. The deep learni...
Elementary Object Systems (EOSs) are a model in the nets-within-nets (NWNs) paradigm, where tokens in turn can host standard Petri nets. We study the complexity of the reachability problem of EOSs when subjected to non-deterministic token losses. It is known that this problem is equivalent to the coverability problem with no lossiness of conservative EOSs (cEOSs). We precisely characterize cEOS ...
Image restoration is a key task in low-level computer vision that aims to reconstruct high-quality images from degraded inputs. The emergence of Vis...
The rise of imaging techniques such as optical coherence tomography (OCT) and advances in deep learning (DL) have enabled clinicians and researchers...
Reasoning is a hallmark of human intelligence, enabling adaptive decision-making in complex and unfamiliar scenarios. In contrast, machine intellige...
We propose a novel approach for estimating the relative pose between rolling shutter cameras using the intersections of line projections with a sing...
Zero-shot Semantic Segmentation (ZSS) aims to segment both seen and unseen classes using supervision from only seen classes. Beyond adaptation-based...
Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semanti...
Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their...
Understanding the intricate workflows of cataract surgery requires modeling complex interactions between surgical tools, anatomical structures, and ...
Glaucoma is a leading cause of irreversible blindness, but early detection can significantly improve treatment outcomes. Traditional diagnostic meth...
Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment...
Recent progress in vision-language segmentation has significantly advanced grounded visual understanding. However, these models often exhibit halluc...
Large Vision-Language Models (LVLMs) have demonstrated significant advancements in multimodal understanding, yet they are frequently hampered by hal...
Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from thei...
Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-L...
Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling comp...
Visual grounding in text-rich document images is a critical yet underexplored challenge for document intelligence and visual question answering (VQA...
Image tokenization plays a critical role in reducing the computational demands of modeling high-resolution images, significantly improving the effic...
Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their suc...