Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 45,121 to 45,130 of 224,055 articles

CoShadow: Multi-Object Shadow Generation for Image Compositing via Diffusion Model

arXiv
Realistic shadow generation is crucial for achieving seamless image compositing, yet existing methods primarily focus on single-object insertion and often fail to generalize when multiple foreground objects are composited into a background scene. In ... read more 

Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing

arXiv
Multimodal large language models (MLLMs) suffer from pronounced hallucinations in remote sensing visual question-answering (RS-VQA), primarily caused by visual grounding failures in large-scale scenes or misinterpretation of fine-grained small target... read more 

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation... read more 

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation... read more 

Designing UNICORN: a Unified Benchmark for Imaging in Computational Pathology, Radiology, and Natural Language

arXiv
Medical foundation models show promise to learn broadly generalizable features from large, diverse datasets. This could be the base for reliable cross-modality generalization and rapid adaptation to new, task-specific goals, with only a few task-spec... read more 

VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning

arXiv
Large models are increasingly becoming autonomous agents that interact with real-world environments and use external tools to augment their static capabilities. However, most recent progress has focused on text-only large language models, which are l... read more 

NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing

arXiv
Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly challenging and constitutes a critical bottleneck, especially for local ... read more 

Structure-Aware Text Recognition for Ancient Greek Critical Editions

arXiv
Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex layout semantics of historical scholarly texts remains limited. This paper investigates structure-awa... read more 

Toward Early Quality Assessment of Text-to-Image Diffusion Models

arXiv
Recent text-to-image (T2I) diffusion and flow-matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate--then--select'' mode: many seeds are sampled and only a... read more 

Toward Early Quality Assessment of Text-to-Image Diffusion Models

arXiv
Recent text-to-image (T2I) diffusion and flow-matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate--then--select'' mode: many seeds are sampled and only a... read more