Latest AI and machine learning research in ophthalmology for healthcare professionals.
The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios. However, existing methods typically feed LLMs with tokens from multi-view and multi-frame images independently, leading to redundant computation and l...
Accurate predictions of the interactions (covalent bonds and non-covalent contacts between atoms) in a molecular system require scalable, accurate, and interpretable energy functions. While classical force fields and knowledge-based energy functions struggle to capture key electronic effects, quantum chemistry approaches such as density functional theory (DFT) provide the necessary accuracy but re...
Vision transformers have demonstrated remarkable success in classification by leveraging global self-attention to capture long-range dependencies. How...
Unlike conventional single-image models, differential medical VQA frameworks process multiple images to identify differences, mirroring the comparativ...
Mobile photography is often limited by complex, lens-specific optical aberrations. While recent deep learning methods approach this as an end-to-end d...
Source Free Unsupervised Domain Adaptation (SFUDA) is critical for deploying deep learning models across diverse clinical settings. However, existing ...
Infrared small target detection (ISTD) is challenging because tiny, low-contrast targets are easily obscured by complex and dynamic backgrounds. Conve...
Astronomers have acquired vast repositories of multimodal data, including images, spectra, and time series, complemented by decades of literature that...
Multimodal Large Language Models (MLLMs) have shown strong performance in vision-language tasks, but their inference efficiency is severely limited by...
Glass surface segmentation from RGB images is a challenging task, since glass as a transparent material distinctly lacks visual characteristics. Howev...
The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting significant attention due to its efficient pro...
When visual evidence is ambiguous, vision models must decide whether to interpret face-like patterns as meaningful. Face pareidolia, the perception of...
Diabetic Retinopathy (DR) requires timely screening to prevent irreversible vision loss. However, its early detection remains a significant challenge ...
Density aggregation is a central problem in machine learning, for instance when combining predictions from a Deep Ensemble. The choice of aggregation ...
Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, vi...
Early screening for glaucoma and diabetic retinopathy (DR) is critical to prevent irreversible vision loss, yet remains inaccessible to many underserv...
Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques ...
The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space ...
Digital advertising increasingly relies on visual content, yet marketers lack rigorous methods for understanding how specific visual attributes causal...
Background: The visual field (VF) test results of many eyes with glaucoma progress despite treatment. This suggests that some eyes are either untreate...