Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 19,371 to 19,380 of 214,800 articles

L2P: Unlocking Latent Potential for Pixel Generation

arXiv
Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohibitive computational and data resources. To address this, we propose the Latent-to-Pixel (L2P) tran... read more 

FAME: Feature Activation Map Explanation on Image Classification and Face Recognition

arXiv
Deep Learning has revolutionized machine learning, reaching unprecedented levels of accuracy, but at the cost of reduced interpretability. Especially in image processing systems, deep networks transform local pixel information into more global concep... read more 

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization

arXiv
Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, are often considerably more complex than image classification, which pr... read more 

Spectral Vision Transformer for Efficient Tokenization with Limited Data

arXiv
We propose a novel spectral vision transformer architecture for efficient tokenization in limited data, with an emphasis on medical imaging. We outline convenient theoretical properties arising from the choice of basis including spatial invariance an... read more 

Resilient Vision-Tabular Multimodal Learning under Modality Missingness

arXiv
Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete modality avai... read more 

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

arXiv
Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric deformations, because the model... read more 

Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection

arXiv
Zero-shot anomaly detection aims to identify defects in unseen categories without target-specific training. Existing methods usually apply the same feature transformation to all samples, treating normal and anomalous data uniformly despite their fund... read more 

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

arXiv
Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physical processes remains difficult, as models must combine object localiza... read more 

The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragments

arXiv
Jigsaw puzzle solving has been an increasingly popular task in the computer vision research community. Recent works have utilized cutting-edge architectures and computational approaches to reassemble groups of pieces into a coherent image, while achi... read more 

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

arXiv
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning:... read more