Latest AI and machine learning research in ophthalmology for healthcare professionals.
Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human activities through video and audio, enabling many human-computer interaction and human-augmentation applications such as human activity support, real-world agents, and skill tra...
We propose a novel approach that adapts hierarchical vision foundation models for real-time ultrasound image segmentation. Existing ultrasound segmentation methods often struggle with adaptability to new tasks, relying on costly manual annotations, while real-time approaches generally fail to match state-of-the-art performance. To overcome these limitations, we introduce an adaptive framework th...
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success in cross-modal tasks such as zero-shot image classification and text-i...
Large vision and language models show strong performance in tasks like image captioning, visual question answering, and retrieval. However, challeng...
The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In ...
Eye-tracking is a vital technology for human-computer interaction, especially in wearable devices such as AR, VR, and XR. The realization of high-sp...
Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly either assess text-...
Medical image segmentation typically relies solely on visual data, overlooking the rich textual information clinicians use for diagnosis. Vision-lan...
Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve L...
We present a framework for optimizing prompts in vision-language models to elicit multimodal reasoning without model retraining. Using an evolutiona...
Referring Remote Sensing Image Segmentation (RRSIS) is a challenging task, aiming to segment specific target objects in remote sensing (RS) images b...
Conversational recommender systems engage users in dialogues to refine their needs and provide more personalized suggestions. Although textual infor...
Vision network designs, including Convolutional Neural Networks and Vision Transformers, have significantly advanced the field of computer vision. Y...
Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessi...
The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential...
We tackle the challenge brought to urban library systems by the {holds system} -- which allows users to request books available at other branches to...
Large Vision-Language Models (LVLMs) struggle with puzzles, which require precise perception, rule comprehension, and logical reasoning. Assessing a...
Shape and texture recognition is fundamental to visual perception. The ability to identify shapes regardless of orientation, texture, or context, an...
The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of...
Vision-based tactile sensors use structured light to measure deformation in their elastomeric interface. Until now, vision-based tactile sensors suc...