Latest AI and machine learning research in ophthalmology for healthcare professionals.
Text-to-image retrieval is a fundamental task in vision-language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries are typically only one or two words long, rendering them semantically ambiguous, prone to collisions across diverse visual interpretations, and lacking explicit control over the quality of retrieved images. To address t...
While Vision-Language Models (VLMs) exhibit exceptional 2D visual understanding, their ability to comprehend and reason about 3D space--a cornerstone of spatial intelligence--remains superficial. Current methodologies attempt to bridge this domain gap either by relying on explicit 3D modalities or by augmenting VLMs with partial, view-conditioned geometric priors. However, such approaches hinder s...
Visual reinforcement learning is appealing for robotics but expensive -- off-policy methods are sample-efficient yet slow; on-policy methods paralleli...
Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, curre...
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry a...
Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders oft...
Unsupervised Anomaly Detection (UAD) aims to identify abnormal regions by establishing correspondences between test images and normal templates. Exist...
Vision language models (VLMs) have achieved remarkable success in broad visual understanding, yet they remain challenged by object-centric reasoning o...
Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, curre...
Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibi...
Vision Transformers have achieved state-of-the-art performance in a wide range of computer vision tasks, but their practical deployment is limited b...
Accurate heat-demand maps play a crucial role in decarbonizing space heating, yet most municipalities lack detailed building-level data needed to calc...
Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We exte...
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry a...
The fine-grained classification of street trees is a crucial task for urban planning, streetscape management, and the assessment of urban ecosystem se...
With the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific informati...
Autonomous navigation in complex, unstructured outdoor environments requires robots to operate over long ranges without prior maps and limited depth s...
Early and accurate classification of retinal diseases is critical to counter vision loss and for guiding clinical management of retinal diseases. In t...
The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent str...
Vision Transformers rely on positional embeddings and class tokens that encode fixed spatial priors. While effective for natural images, these priors ...