Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 47,271 to 47,280 of 224,199 articles

Therapist-Robot-Patient Physical Interaction is Worth a Thousand Words: Enabling Intuitive Therapist Guidance via Remote Haptic Control

arXiv
Robotic systems can enhance the amount and repeatability of physically guided motor training. Yet their real-world adoption is limited, partly due to non-intuitive trainer/therapist-trainee/patient interactions. To address this gap, we present a hapt... read more 

SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model

arXiv
SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branch synthesizes video and the ot... read more 

SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model

arXiv
SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branch synthesizes video and the ot... read more 

SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance

arXiv
Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending this succe... read more 

Joint Shadow Generation and Relighting via Light-Geometry Interaction Maps

arXiv
We propose Light-Geometry Interaction (LGI) maps, a novel representation that encodes light-aware occlusion from monocular depth. Unlike ray tracing, which requires full 3D reconstruction, LGI captures essential light-shadow interactions reliably and... read more 

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

arXiv
Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce StoryMovie, a dataset of 1,757 stor... read more 

UniVBench: Towards Unified Evaluation for Video Foundation Models

arXiv
Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation benchmarks re... read more 

Meta-FC: Meta-Learning with Feature Consistency for Robust and Generalizable Watermarking

arXiv
Deep learning-based watermarking has made remarkable progress in recent years. To achieve robustness against various distortions, current methods commonly adopt a training strategy where a \underline{\textbf{s}}ingle \underline{\textbf{r}}andom \unde... read more 

DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs

arXiv
Vision-Language Models (VLMs) have emerged as versatile solutions for zero-shot question answering (QA) across various domains. However, enabling VLMs to effectively comprehend structured graphs and perform accurate, efficient QA remains challenging.... read more 

GFPL: Generative Federated Prototype Learning for Resource-Constrained and Data-Imbalanced Vision Task

arXiv
Federated learning (FL) facilitates the secure utilization of decentralized images, advancing applications in medical image recognition and autonomous driving. However, conventional FL faces two critical challenges in real-world deployment: ineffecti... read more