Latest AI and machine learning research in alternative medicine for healthcare professionals.
4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling embodied interaction. While current Multimodal Large Language Models (MLLMs) show strong capabilities in static scene understanding and coarse-grained 4D tasks, they still have notable limitations in continuous dynamic scene perception, especially i...
Freehand 3-D ultrasound (US) imaging has attracted increasing attention owing to its intuitive volumetric visualization, ease of use, and low cost. However, accurate 3-D reconstruction critically depends on stable probe pose estimation, yet existing trackerless methods remain susceptible to accumulated pose errors, particularly over long scanning trajectories. To address this limitation, we propos...
4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling...
Multimodal anomaly detection benefits from complementary RGB and 3D evidence, yet auxiliary RGB reconstruction is not equally reliable across product ...
Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection. Recen...
Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retriev...
Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial sca...
Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity. While data synthes...
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are ...
Allosteric regulation represents a fundamental mechanism of protein function, yet distinguishing allosteric from orthosteric protein binding sites rem...
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact...
Stroke is a leading cause of death and long-term disability worldwide, affecting approximately 15 million individuals annually. Prompt and accurate su...
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens...
In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language...
Agentic large language models are increasingly used across the genomic workflow, from variant calling to clinical interpretation, yet they are evaluat...
Chromosomal copy number alterations (CNAs) are key drivers of tumor evolution, disease progression and therapeutic resistance, and the identification ...
Accurate prediction of stress and strain fields in hierarchical composite microstructures is critical for physics-informed material design, yet conven...
The proliferation of multimedia content on social platforms has fueled multimodal misinformation, where images are used to reinforce false claims. Con...
Objectives: To characterize residual false positives in prostate MRI detection, and to evaluate a lightweight post-hoc refinement head for case-level ...
Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpre...