Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 45,171 to 45,180 of 224,055 articles

mHC-HSI: Clustering-Guided Hyper-Connection Mamba for Hyperspectral Image Classification

arXiv
Recently, DeepSeek has invented the manifold-constrained hyper-connection (mHC) approach which has demonstrated significant improvements over the traditional residual connection in deep learning models \cite{xie2026mhc}. Nevertheless, this approach h... read more 

Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning

arXiv
Recent work shows that text-only reinforcement learning with verifiable rewards (RLVR) can match or outperform image-text RLVR on multimodal medical VQA benchmarks, suggesting current evaluation protocols may fail to measure causal visual dependence.... read more 

Biased Generalization in Diffusion Models

arXiv
Generalization in generative modeling is defined as the ability to learn an underlying distribution from a finite dataset and produce novel samples, with evaluation largely driven by held-out performance and perceived sample quality. In practice, tra... read more 

Beyond Pixel Histories: World Models with Persistent 3D State

arXiv
Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D representation of the environment, meaning 3D consistency must be implici... read more 

Optimal trajectory-guided stochastic co-optimization for e-fuel system design and real-time operation

arXiv
E-fuels are promising long-term energy carriers supporting the net-zero transition. However, the large combinatorial design-operation spaces under renewable uncertainty make the use of mathematical programming impractical for co-optimizing e-fuel pro... read more 

Logit-Level Uncertainty Quantification in Vision-Language Models for Histopathology Image Analysis

arXiv
Vision-Language Models (VLMs) with their multimodal capabilities have demonstrated remarkable success in almost all domains, including education, transportation, healthcare, energy, finance, law, and retail. Nevertheless, the utilization of VLMs in h... read more 

PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest

arXiv
While multi-modal Visual Language Models (VLMs) have demonstrated significant success across various domains, the integration of VLMs into recommendation and retrieval systems remains a challenge, due to issues like training objective discrepancies a... read more 

Learning functional groups in complex microbiomes

arXiv
From soil to the gut, communities composed of thousands of microbes perform functions such as carbon sequestration and immune system regulation. Here, we introduce a data-driven approach that explains how community function can be traced to just a fe... read more 

Modeling Cross-vision Synergy for Unified Large Vision Model

arXiv
Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, videos, and 3D data. However, existing unified LVMs primarily pursue functional integration, while ove... read more 

Confidence-aware Monocular Depth Estimation for Minimally Invasive Surgery

arXiv
Purpose: Monocular depth estimation (MDE) is vital for scene understanding in minimally invasive surgery (MIS). However, endoscopic video sequences are often contaminated by smoke, specular reflections, blur, and occlusions, limiting the accuracy of ... read more