Latest AI and machine learning research in ophthalmology for healthcare professionals.
Agricultural domains are being transformed by recent advances in AI and computer vision that support quantitative visual evaluation. Using aerial and ground imaging over a time series, we develop a framework for characterizing the ripening process of cranberry crops, a crucial component for precision agriculture tasks such as comparing crop breeds (high-throughput phenotyping) and detecting dise...
Vision-Language Models (VLMs) have shown promising capabilities in handling various multimodal tasks, yet they struggle in long-context scenarios, particularly in tasks involving videos, high-resolution images, or lengthy image-text documents. In our work, we first conduct an empirical analysis of the long-context capabilities of VLMs using our augmented long-context multimodal datasets. Our fin...
Large Vision-Language Models (VLMs) have been extended to understand both images and videos. Visual token compression is leveraged to reduce the con...
Existing multi-modal learning methods on fundus and OCT images mostly require both modalities to be available and strictly paired for training and t...
Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language...
Human readers can accurately count how many letters are in a word (e.g., 7 in ``buffalo''), remove a letter from a given position (e.g., ``bufflo'')...
Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs...
Vision-language models have made significant strides recently, demonstrating superior performance across a range of tasks, e.g. optical character re...
Visible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods ...
Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usu...
Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common ...
The high-performance computing (HPC) landscape is undergoing rapid transformation, with an increasing emphasis on energy-efficient and heterogeneous...
DALL-E and Sora have gained attention by producing implausible images, such as "astronauts riding a horse in space." Despite the proliferation of te...
Flare, an optical phenomenon resulting from unwanted scattering and reflections within a lens system, presents a significant challenge in imaging. T...
Textured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D mode...
Despite the efficiency of prompt learning in transferring vision-language models (VLMs) to downstream tasks, existing methods mainly learn the promp...
In recent years, Visual Question Answering (VQA) has made significant strides, particularly with the advent of multimodal models that integrate visi...
This paper provides a comprehensive review of mechanical equipment fault diagnosis methods, focusing on the advancements brought by Transformer-base...
Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high...
We propose a method for dense depth estimation from an event stream generated when sweeping the focal plane of the driving lens attached to an event...