Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 33,281 to 33,290 of 221,422 articles

FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On

arXiv
Given a person and a garment image, virtual try-on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their original pose and identity. Although recent VTO methods excel at visualizing garment appearance, t... read more 

RewardFlow: Generate Images by Optimizing What You Reward

arXiv
We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward Langevin dynamics. RewardFlow unifies complementary differentiable rewards for semantic alignment, p... read more 

Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding

arXiv
Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. A field-wide goal is to achieve generalizable, cro... read more 

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts

arXiv
Multimodal Mixture-of-Experts (MoE) models have achieved remarkable performance on vision-language tasks. However, we identify a puzzling phenomenon termed Seeing but Not Thinking: models accurately perceive image content yet fail in subsequent reaso... read more 

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

arXiv
This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D priors or ge... read more 

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

arXiv
The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a profound meta-cognitive deficit: they struggle to arbitrate between leveraging internal knowledge and... read more 

On Semiotic-Grounded Interpretive Evaluation of Generative Art

arXiv
Interpretation is essential to deciphering the language of art: audiences communicate with artists by recovering meaning from visual artifacts. However, current Generative Art (GenArt) evaluators remain fixated on surface-level image quality or liter... read more 

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation

arXiv
Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same time, high-quality video editing data remains scarce. In this paper, we... read more 

EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition

arXiv
How do you build a sign language recognizer that works on a phone? That question drove this work. We built EfficientSign, a lightweight model which takes EfficientNet-B0 and focuses on two attention modules (Squeeze-and-Excitation for channel focus, ... read more 

EvoLen: Evolution-Guided Tokenization for DNA Language Model

arXiv
Tokens serve as the basic units of representation in DNA language models (DNALMs), yet their design remains underexplored. Unlike natural language, DNA lacks inherent token boundaries or predefined compositional rules, making tokenization a fundament... read more