Latest AI and machine learning research in ophthalmology for healthcare professionals.
Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propo...
Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through...
Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality imag...
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handle...
The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to transform myopi...
The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing thi...
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features ...
Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In cl...
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundam...
Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in...
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstructio...
Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morpholog...
Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pre...
End-to-end OCR systems based on vision-language models have achieved strong performance in complex document OCR, but their efficiency is limited by th...
While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmar...
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units enc...
Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic und...
Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. Ho...
Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reduci...
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiri...