Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 44,361 to 44,370 of 224,055 articles

FontUse: A Data-Centric Approach to Style- and Use-Case-Conditioned In-Image Typography

arXiv
Recent text-to-image models can generate high-quality images from natural-language prompts, yet controlling typography remains challenging: requested typographic appearance is often ignored or only weakly followed. We address this limitation with a d... read more 

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

arXiv
Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a persistent capabil... read more 

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

arXiv
The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long tail scenarios. However, these models often fail on ... read more 

DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model

arXiv
Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in visual dat... read more 

Cross-Resolution Distribution Matching for Diffusion Distillation

arXiv
Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising process, where step reduction has largely saturated. Partial timestep low-resolution generation can further ... read more 

Spatial Colour Mixing Illusions as a Perception Stress Test for Vision-Language Models

arXiv
Vision-language models (VLMs) achieve strong benchmark results, yet can exhibit systematic perceptual weaknesses: structured, large changes to pixel values can cause confident yet nonsensical predictions, even when the underlying scene remains easily... read more 

Longitudinal NSCLC Treatment Progression via Multimodal Generative Models

arXiv
Predicting tumor evolution during radiotherapy is a clinically critical challenge, particularly when longitudinal changes are driven by both anatomy and treatment. In this work, we introduce a Virtual Treatment (VT) framework that formulates non-smal... read more 

VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models

arXiv
Vision-language models (VLMs) achieve strong performance on standard, high-quality datasets, but we still do not fully understand how they perform under real-world image distortions. We present VLM-RobustBench, a benchmark spanning 49 augmentation ty... read more 

Reflective Flow Sampling Enhancement

arXiv
The growing demand for text-to-image generation has led to rapid advances in generative modeling. Recently, text-to-image diffusion models trained with flow matching algorithms, such as FLUX, have achieved remarkable progress and emerged as strong al... read more 

Wearable Sleep Measures May Improve Machine Learning Prediction of Home-Based Pulmonary Rehabilitation Engagement Among Patients With Chronic Obstructive Pulmonary Disease: A Proof-of-Concept Study.

Mayo Clinic proceedings. Digital health
OBJECTIVE: To evaluate whether incorporating baseline sleep measures from a wrist-worn activity monitor in machine learning (ML) models improved the prediction of 12-week engagement with home-based pulmonary rehabilitation (HBPR) in patients with chr... read more