Latest AI and machine learning research in ophthalmology for healthcare professionals.
Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by textual queries. However, widely adopted naive fine-tuning strategies concentrate exclusively on text-image consistency for categories contained in training, which leads to limited generalizability for unseen categorie...
We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for environmental interaction and collaboration with autonomous agents. Despite advancements in spatial r...
We present DyMU, an efficient, training-free framework that dynamically reduces the computational burden of vision-language models (VLMs) while main...
Image description generation is essential for accessibility and AI understanding of visual content. Recent advancements in deep learning have signif...
Streetscapes are an essential component of urban space. Their assessment is presently either limited to morphometric properties of their mass skelet...
In recent years, large-scale vision-language models (VLMs) like CLIP have gained attention for their zero-shot inference using instructional text pr...
Automated breast cancer detection via computer vision techniques is challenging due to the complex nature of breast tissue, the subtle appearance of...
Artificial intelligence (AI) shows remarkable potential in medical imaging diagnostics, yet most current models require retraining when applied acro...
Diabetic retinopathy is a serious ocular complication that poses a significant threat to patients' vision and overall health. Early detection and ac...
We operate through the lens of ordinary differential equations and control theory to study the concept of observability in the context of neural sta...
Preference alignment through Direct Preference Optimization (DPO) has demonstrated significant effectiveness in aligning multimodal large language m...
Recognizing and reasoning about occluded (partially or fully hidden) objects is vital to understanding visual scenes, as occlusions frequently occur...
Distinguishing between real and AI-generated images, commonly referred to as 'image detection', presents a timely and significant challenge. Despite...
Eye-tracking analysis plays a vital role in medical imaging, providing key insights into how radiologists visually interpret and diagnose clinical c...
Vision Transformers (ViTs) have revolutionized computer vision by leveraging self-attention to model long-range dependencies. However, ViTs face cha...
A driver face monitoring system can detect driver fatigue, which is a significant factor in many accidents, using computer vision techniques. In thi...
The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modaliti...
Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fal...
Vision Large Language Models (VLLMs) have demonstrated impressive capabilities in general visual tasks such as image captioning and visual question ...
Zero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of a...