Latest AI and machine learning research in ophthalmology for healthcare professionals.
The application of Vision-Language Models (VLMs) in remote sensing (RS) has demonstrated significant potential in traditional tasks such as scene classification, object detection, and image captioning. However, current models, which excel in Referring Expression Comprehension (REC), struggle with tasks involving complex instructions (e.g., exists multiple conditions) or pixel-level operations li...
Encoder-free multimodal large language models(MLLMs) eliminate the need for a well-trained vision encoder by directly processing image tokens before the language model. While this approach reduces computational overhead and model complexity, it often requires large amounts of training data to effectively capture the visual knowledge typically encoded by vision models like CLIP. The absence of a ...
We propose the notion of empirical privacy variance and study it in the context of differentially private fine-tuning of language models. Specifical...
Deciphering the neural mechanisms that transform sensory experiences into meaningful semantic representations is a fundamental challenge in cognitiv...
Anterior Segment Optical Coherence Tomography (AS-OCT) is an emerging imaging technique with great potential for diagnosing anterior uveitis, a visi...
Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current eff...
Text-to-Image (T2I) Diffusion Models have achieved remarkable performance in generating high quality images. However, enabling precise control of co...
Vision-Language Models (VLMs) leverage aligned visual encoders to transform images into visual tokens, allowing them to be processed similarly to te...
Deploying large vision-language models (LVLMs) introduces a unique vulnerability: susceptibility to malicious attacks via visual inputs. However, ex...
Current Cross-Modality Generation Models (GMs) demonstrate remarkable capabilities in various generative tasks. Given the ubiquity and information r...
Language-guided attention frameworks have significantly enhanced both interpretability and performance in image classification; however, the relianc...
Visual Question Answering (VQA) models, which fall under the category of vision-language models, conventionally execute multiple downsampling proces...
Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following huma...
Vision and Language Navigation (VLN) requires an agent to navigate through environments following natural language instructions. However, existing m...
BACKGROUND AND HYPOTHESIS: Substantive inquiry into the predictive power of eye movement (EM) features for clinical high-risk (CHR) conversion and the...
Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities b...
Roadside vision centric 3D object detection has received increasing attention in recent years. It expands the perception range of autonomous vehicle...
Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture fo...
We propose a general framework called VisionLogic to extract interpretable logic rules from deep vision models, with a focus on image classification...
Learning from noisy ordinal labels is a key challenge in medical imaging. In this work, we ask whether ordinal disease progression labels (better, w...