Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 38,251 to 38,260 of 223,469 articles

Longitudinal Digital Phenotyping for Early Cognitive-Motor Screening

arXiv
Early detection of atypical cognitive-motor development is critical for timely intervention, yet traditional assessments rely heavily on subjective, static evaluations. The integration of digital devices offers an opportunity for continuous, objectiv... read more 

Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming

arXiv
Cross-view geo-localization (CVGL) estimates a camera's location by matching a street-view image to geo-referenced overhead imagery, enabling GPS-denied localization and navigation. Existing methods almost universally formulate CVGL as an image-retri... read more 

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

arXiv
Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of producing interleaved content i... read more 

TRACE: Object Motion Editing in Videos with First-Frame Trajectory Guidance

arXiv
We study object motion path editing in videos, where the goal is to alter a target object's trajectory while preserving the original scene content. Unlike prior video editing methods that primarily manipulate appearance or rely on point-track-based t... read more 

No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models

arXiv
Contrastive vision-language (V&L) models remain a popular choice for various applications. However, several limitations have emerged, most notably the limited ability of V&L models to learn compositional representations. Prior methods often addressed... read more 

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

arXiv
We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation from both RGB-only and RGB-D inputs. While recent works with foundation approaches have shown that an increase in the quantity and... read more 

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

arXiv
Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly focus on natural image synthesis and fail to systematically evaluate mo... read more 

PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow

arXiv
Graphic design is a creative and innovative process that plays a crucial role in applications such as e-commerce and advertising. However, developing an automated design system that can faithfully translate user intentions into editable design files ... read more 

RefAlign: Representation Alignment for Reference-to-Video Generation

arXiv
Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice... read more 

MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models

arXiv
Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow these models to handle varying input sizes during training, inference typ... read more