Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 18,711 to 18,720 of 214,544 articles

MonoPRIO: Adaptive Prior Conditioning for Unified Monocular 3D Object Detection

arXiv
Monocular 3D object detection remains challenging because metric size and depth are underdetermined by single-view evidence, particularly under occlusion, truncation, and projection-induced scale-depth ambiguity. Although recent methods improve depth... read more 

Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

arXiv
Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR benchmarks is as... read more 

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

arXiv
In recent years, computer vision has witnessed remarkable progress, fueled by the development of innovative architectures such as Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), diffusion-based architectures, Vision Tran... read more 

The Velocity Deficit: Initial Energy Injection for Flow Matching

arXiv
While Flow Matching theoretically guarantees constant-velocity trajectories, we identify a critical breakdown in high-dimensional practice: the Velocity Deficit. We show that the MSE objective systematically underestimates velocity magnitude, causing... read more 

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

arXiv
Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit from strong generative priors, most methods still condition only on low-quality inputs, making it... read more 

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

arXiv
Humans naturally communicate through abstract concepts like "mood". However, current image editing benchmarks focus primarily on explicit, literal commands, leaving abstract instructions largely underexplored. In this work, we first formalize the def... read more 

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

arXiv
Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion governed by kinematic and geometric constraints. In these settings, object parts must remain rigid... read more 

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study

arXiv
Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigorous biometric tasks remains unexplored. This work presents an exploratory study evaluating the zer... read more 

SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

arXiv
Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and texture distortions that degrade perceived quality. These defects vary widely in perceptual impact-... read more 

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

arXiv
Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively wel... read more