Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 39,421 to 39,430 of 223,737 articles

SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning

arXiv
Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are predominantly trained on 2D image data and therefore often fail to capture 3D spatial relationships be... read more 

P-Flow: Prompting Visual Effects Generation

arXiv
Recent advancements in video generation models have significantly improved their ability to follow text prompts. However, the customization of dynamic visual effects, defined as temporally evolving and appearance-driven visual phenomena like object c... read more 

Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding

arXiv
Text-driven video moment retrieval (VMR) remains challenging due to limited capture of hidden temporal dynamics in untrimmed videos, leading to imprecise grounding in long sequences. Traditional methods rely on natural language queries (NLQs) or stat... read more 

Biophysics-Enhanced Neural Representations for Patient-Specific Respiratory Motion Modeling

arXiv
A precise spatial delivery of the radiation dose is crucial for the treatment success in radiotherapy. In the lung and upper abdominal region, respiratory motion introduces significant treatment uncertainties, requiring special motion management tech... read more 

DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment

arXiv
Reducing token count is crucial for efficient training and inference of latent diffusion models, especially at high resolution. A common strategy is to build high-compression image tokenizers with more channels per token. However, when trained only f... read more 

Multimodal Survival Analysis with Locally Deployable Large Language Models

arXiv
We study multimodal survival analysis integrating clinical text, tabular covariates, and genomic profiles using locally deployable large language models (LLMs). As many institutions face tight computational and privacy constraints, this setting motiv... read more 

Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement

arXiv
Recent advances in Multimodal Large Language Models (MLLMs) have enabled automated generation of structured layouts from natural language descriptions. Existing methods typically follow a code-only paradigm that generates code to represent layouts, w... read more 

PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation

arXiv
Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis that pred... read more 

CayleyPy-4: AI-Holography. Towards analogs of holographic string dualities for AI tasks

arXiv
This is the fourth paper in the CayleyPy project, which applies AI methods to the exploration of large graphs. In this work, we suggest the existence of a new discrete version of holographic string dualities for this setup, and discuss their relevanc... read more 

Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance Learning

arXiv
Multiple Instance Learning (MIL) is the predominant framework for classifying gigapixel whole-slide images in computational pathology. MIL follows a sequence of 1) extracting patch features, 2) applying a linear layer to obtain task-specific patch fe... read more