Latest AI and machine learning research in medicare for healthcare professionals.
Weakly supervised object localization (WSOL) aims to localize target objects in images using only image-level labels. Despite recent progress, many approaches still rely on multi-stage pipelines or full fine-tuning of large backbones, which increases training cost, while the broader WSOL community continues to face the challenge of partial object coverage. We present TriLite, a single-stage WSOL f...
Real-world multimodal agents solve multi-step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip by interpreting a transit map and checking schedules under routing constraints. However, existing multimodal benchmarks mainly evaluate single-turn visual reason...
Quantifying uncertainty in clinical predictions is critical for high-stakes diagnosis tasks. Conformal prediction offers a principled approach by prov...
High genomic variability among viral species makes sequence classification highly dependent on multiple sequence alignment (MSA) methods, which are bo...
Single-image 3D generation with part-level structure remains challenging: learned priors struggle to cover the long tail of part geometries and mainta...
Medical vision-language models (VLMs) are strong zero-shot recognizers for medical imaging, but their reliability under domain shift hinges on calibra...
Face morphing attacks are widely recognized as one of the most challenging threats to face recognition systems used in electronic identity documents. ...
Background and Aims: Alcohol use disorder (AUD) remains a major public health concern, with persistent disparities in access to evidence-based treatme...
We propose a novel method for establishing correspondence between two sequences of 2D images. One particular application of this technique is slice-le...
We present MedXIAOHE, a medical vision-language foundation model designed to advance general-purpose medical understanding and reasoning in real-world...
Fanconi anemia (FA) is a rare genetic disorder of impaired DNA repair characterized by progressive bone marrow failure, congenital malformations, and ...
Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This p...
Accurate prediction of outcomes is crucial for clinical decision-making and personalized patient care. Supervised machine learning algorithms, which a...
We present a principled framework for confidence estimation in computed tomography (CT) reconstruction. Based on the sequential likelihood mixing fram...
Large vision-language models such as CLIP struggle with long captions because they align images and texts as undifferentiated wholes. Fine-grained vis...
Large Language Models (LLMs) trained for average correctness often exhibit mode collapse, producing narrow decision behaviors on tasks where multiple ...
Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal co...
Prediction sets can wrap around any ML model to cover unknown test outcomes with a guaranteed probability. Yet, it remains unclear how to use them opt...
Background Patients with repaired tetralogy of Fallot (rTOF) require lifelong surveillance with cardiovascular magnetic resonance (CMR) and cardiopulm...
Graph neural networks (GNNs) have become the standard tool for encoding data and their complex relationships into continuous representations, improvin...