Latest AI and machine learning research in alternative medicine for healthcare professionals.
Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study investigates an explainable DR classification framework using vision foundation models and multiple transfer learning strategies. Three backbones, DINOv2, CLIP, and Vision Transformer (ViT), were evaluated using full fine-tuning, linear probing, and Low-...
Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. In this randomized mixed-methods study, 34 oncology physicians provided 294 ratings of four blinded AI systems across five vignettes, alongside 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar refer...
Medical vision-language models (VLMs) can appear reliable in-domain while failing when acquisition domain, paired supervision, or evaluation protocol ...
Dataset distillation compresses a large training set into a compact synthetic set while preserving its downstream utility. However, existing methods p...
Volumetric medical VQA requires reasoning over long and redundant 3D visual token sequences, especially in multi-sequence MRI where complementary moda...
Multimodal product models can complete missing e-commerce attributes, yet current methods still optimize attribute-answer accuracy without verifying v...
Background: Many of the most consequential treatment decisions concern patients and comparisons that randomized trials never address: off-label and he...
Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evi...
Multi-teacher distillation has emerged as a way to combine complementary teacher models into a single student model that exhibits the strengths of all...
Spatial multi-omics technologies jointly profile gene expression, surface proteins, and histology at each tissue spot, yet most spatial domain discove...
The interpretation of endoscopic imagery in ulcerative colitis is complex and subjective, with variability in human assessment and subtle mucosal infl...
Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrati...
Background: Clinical terminology pipelines must first extract candidate spans from narrative notes and then determine whether those spans map to exist...
Spatial multi-omics technologies jointly profile gene expression, surface proteins, and histology at each tissue spot, yet most spatial domain discove...
Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Exis...
Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justi...
Missense mutations and post-translational modifications (PTMs) are major molecular perturbations that reshape protein function but are traditionally s...
Trustworthy multimodal fusion in clinical settings requires handling incomplete and heterogeneous modality subsets across institutions, where privacy ...
Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual e...
Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they recei...