Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 40,571 to 40,580 of 223,737 articles

Harnessing the Power of Foundation Models for Accurate Material Classification

arXiv
Material classification has emerged as a critical task in computer vision and graphics, supporting the assignment of accurate material properties to a wide range of digital and real-world applications. While traditionally framed as an image classific... read more 

Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation

arXiv
Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture labels is available and show ... read more 

Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion

arXiv
We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content uniformly, our method dynamically adjusts temporal attention receptiv... read more 

Joint Degradation-Aware Arbitrary-Scale Super-Resolution for Variable-Rate Extreme Image Compression

arXiv
Recent diffusion-based extreme image compression methods have demonstrated remarkable performance at ultra-low bitrates. However, most approaches require training separate diffusion models for each target bitrate, resulting in substantial computation... read more 

Towards Motion-aware Referring Image Segmentation

arXiv
Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion-related queries compared to appearance-based ones. To address this, we fi... read more 

Structured SIR: Efficient and Expressive Importance-Weighted Inference for High-Dimensional Image Registration

arXiv
Image registration is an ill-posed dense vision task, where multiple solutions achieve similar loss values, motivating probabilistic inference. Variational inference has previously been employed to capture these distributions, however restrictive ass... read more 

SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning

arXiv
Image-conditioned Video diffusion models achieve impressive visual realism but often suffer from weakened motion fidelity, e.g., reduced motion dynamics or degraded long-term temporal coherence, especially after fine-tuning. We study the problem of m... read more 

AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement

arXiv
GUI grounding is a critical capability for vision-language models (VLMs) that enables automated interaction with graphical user interfaces by locating target elements from natural language instructions. However, grounding on GUI screenshots remains c... read more 

UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models

arXiv
Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting... read more 

Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation

arXiv
We present Omni-I2C, a comprehensive benchmark designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executable code. We argue that this task represents a non-trivial challenge f... read more