Latest AI and machine learning research in alternative medicine for healthcare professionals.
Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images - where critical evidence is tiny, subtle, spatially distant, or distributed - remains unclear. Existing evaluations largely report final-answer accuracy, offering limited insight into whether models acquire and integrate the necessary visual evidenc...
While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mitigates this by formulating predictions as a Dirichlet distribution over class probabilities to explicitly quantify epistemic uncertainty. However, we found that the conventional EDL suffers from two fundamental limitations: a Kullback-Leibler (KL)...
Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human exper...
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such...
Benchmarks increasingly guide deployment, procurement and scientific screening, yet a score supports only the response it records, not necessarily the...
Clinical AI systems have achieved strong predictive performance; however, prediction accuracy is not sufficient for clinical safety. Retrieval-augment...
Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential...
Background: African Americans (AA) experience disproportionate burden of colorectal cancer (CRC). Dysregulation of the Wingless-related integration si...
Accurate risk stratification of pigmented skin lesions is critical for early melanoma detection and for reducing unnecessary excisions. Artificial int...
Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by...
Contrastively trained vision-language models such as CLIP provide strong zero-shot transfer by aligning images and text in a shared embedding space. H...
Manual Pap smear analysis for cervical cancer screening is limited by inter-observer variability, time constraints, and restricted expert availability...
Learned image compression (LIC) increasingly requires reconstructions that balance distortion fidelity and perceptual realism across a wide range of b...
Monocular 3D object detection remains challenging because metric size and depth are underdetermined by single-view evidence, particularly under occlus...
Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabi...
Disease screening is critical for early detection and timely intervention in clinical practice. However, most current screening models for medical ima...
Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe sc...
Working memory (WM), the human brain's system for maintaining and manipulating information over short timescales, is critical for goal-directed behavi...
Translational medicine turns underspecified development goals into evidence synthesis that must combine literature, trials, patents, and quantitative ...
The opaque nature of deep learning models remains a significant barrier to their clinical adoption in medical imaging. This paper presents a multimoda...