Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 40,441 to 40,450 of 223,737 articles

ARIADNE: A Perception-Reasoning Synergy Framework for Trustworthy Coronary Angiography Analysis

arXiv
Conventional pixel-wise loss functions fail to enforce topological constraints in coronary vessel segmentation, producing fragmented vascular trees despite high pixel-level accuracy. We present ARIADNE, a two-stage framework coupling preference-align... read more 

Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting

arXiv
Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning for various downstream tasks, such as semantic segmentation, 3D object d... read more 

Tinted Frames: Question Framing Blinds Vision-Language Models

arXiv
Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that VLMs are selectively blind. They modulate the amount of attention appli... read more 

Tinted Frames: Question Framing Blinds Vision-Language Models

arXiv
Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that VLMs are selectively blind. They modulate the amount of attention appli... read more 

RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing

arXiv
Diffusion models have become the dominant paradigm for image generation and editing, with latent diffusion models shifting denoising to a compact latent space for efficiency and scalability. Recent attempts to leverage pretrained visual representatio... read more 

Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders

arXiv
Large vision--language models (VLMs) often use a frozen vision backbone, whose image features are mapped into a large language model through a lightweight connector. While transformer-based encoders are the standard visual backbone, we ask whether st... read more 

DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

arXiv
With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visual modality. However, most existing tokenizers are designed for monocu... read more 

Spectrally-Guided Diffusion Noise Schedules

arXiv
Denoising diffusion models are widely used for high-quality image and video generation. Their performance depends on noise schedules, which define the distribution of noise levels applied during training and the sequence of noise levels traversed dur... read more 

Under One Sun: Multi-Object Generative Perception of Materials and Illumination

arXiv
We introduce Multi-Object Generative Perception (MultiGP), a generative inverse rendering method for stochastic sampling of all radiometric constituents -- reflectance, texture, and illumination -- underlying object appearance from a single image. Ou... read more 

MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction

arXiv
Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and object structu... read more