Latest AI and machine learning research in ophthalmology for healthcare professionals.
Age-related macular degeneration (AMD) is a progressive degenerative retinal disease and a leading cause of blindness in older adults worldwide. According to numerous studies, the number of affected individuals reached 196 million in 2020, with projections estimating an increase to 288 million by 2040, including 18.6 million cases of advanced AMD. The advent of optical coherence tomography (OCT) h...
Despite recent advances in Vision-Language Models (VLMs), they may over-rely on visual language priors existing in their training data rather than true visual reasoning. To investigate this, we introduce ViLP, a benchmark featuring deliberately out-of-distribution images synthesized via image generation models and out-of-distribution Q&A pairs. Each question in ViLP is coupled with three potenti...
Diffusion models have gained tremendous success in text-to-image generation, yet still lag behind with visual understanding tasks, an area dominated...
OBJECTIVES: To assess the appropriateness and readability of large language model (LLM) chatbots' answers to frequently asked questions about refracti...
This paper presents a framework for developing a live vision-correcting display (VCD) to address refractive visual aberrations without the need for ...
A minimalist vision system uses the smallest number of pixels needed to solve a vision task. While traditional cameras use a large grid of square pi...
Transparent objects are ubiquitous in daily life, making their perception and robotics manipulation important. However, they present a major challen...
Large-scale Vision-Language Models (VLMs) have advanced by aligning vision inputs with text, significantly improving performance in computer vision ...
The domain gap between remote sensing imagery and natural images has recently received widespread attention and Vision-Language Models (VLMs) have d...
Recently, "visual o1" began to enter people's vision, with expectations that this slow-thinking design can solve visual reasoning tasks, especially ...
Acquisition and modeling of polarized light reflection and scattering help reveal the shape, structure, and physical characteristics of an object, w...
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance in complex multimodal tasks. However, these models still suffer from h...
Autonomous vehicles face significant challenges in navigating adverse weather, particularly rain, due to the visual impairment of camera-based syste...
Visual Mamba is an approach that extends the selective space state model, Mamba, to vision tasks. It processes image tokens sequentially in a fixed ...
We present a comprehensive theoretical framework analyzing the relationship between data distributions and fairness guarantees in equitable deep lea...
A novel framework is proposed that combines multi-resonance biosensors with machine learning (ML) to significantly enhance the accuracy of parameter...
The autonomous driving community is increasingly focused on addressing corner case problems, particularly those related to ensuring driving safety u...
In the construction sector, workers often endure prolonged periods of high-intensity physical work and prolonged use of tools, resulting in injuries...
The application of Contrastive Language-Image Pre-training (CLIP) in Weakly Supervised Semantic Segmentation (WSSS) research powerful cross-modal se...
Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enh...