Latest AI and machine learning research in ophthalmology for healthcare professionals.
The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based approach to process images, which leads to fragmented visual encoding and a sharp increase in redundant tokens. To tackle these issues, we propose the FALCON model...
Source camera identification has emerged as a vital solution to unlock incidents involving critical cases like terrorism, violence, and other criminal activities. The ability to trace the origin of an image/video can aid law enforcement agencies in gathering evidence and constructing the timeline of events. Moreover, identifying the owner of a certain device narrows down the area of search in a ...
The Visual Language Model, known for its robust cross-modal capabilities, has been extensively applied in various computer vision tasks. In this pap...
Vision-language pretraining (VLP) has been investigated to generalize across diverse downstream tasks for fundus image analysis. Although recent met...
Alzheimer's disease (AD) is a neurodegenerative disorder affecting millions worldwide, necessitating early and accurate diagnosis for optimal patien...
Intelligent vehicular communication with vehicle road collaboration capability is a key technology enabled by 6G, and the integration of various vis...
By overlaying time-synced user comments on videos, Danmu creates a co-watching experience for online viewers. However, its visual-centric design pos...
Existing eye trackers use cameras based on thick compound optical elements, necessitating the cameras to be placed at focusing distance from the eye...
Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researc...
The communication scenarios and channel characteristics of 6G will be more complex and difficult to characterize. Conventional methods for channel p...
This work investigates the capabilities of current vision-language models (VLMs) in visual understanding and attribute measurement of primitive shap...
Vision language model (VLM) has been designed for large scale image-text alignment as a pretrained foundation model. For downstream few shot classif...
Vision systems in nature show remarkable diversity, from simple light-sensitive patches to complex camera eyes with lenses. While natural selection ...
Vision language models have achieved impressive results across various fields. However, adoption in remote sensing remains limited, largely due to t...
This paper introduces an innovative software system for fundus image analysis that deliberately diverges from the conventional screening approach, o...
Automated segmentation plays a pivotal role in medical image analysis and computer-assisted interventions. Despite the promising performance of exis...
This research addresses the challenge of camera calibration and distortion parameter prediction from a single image using deep learning models. The ...
Accessing visual information is crucial yet challenging for people with low vision due to their visual conditions (e.g., low visual acuity, limited ...
Neovascular age-related macular degeneration (nAMD) is a leading cause of vision loss among older adults, where disease activity detection and progr...
As the demand for high-resolution image processing in Large Vision-Language Models (LVLMs) grows, sub-image partitioning has become a popular approa...