Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 20,041 to 20,050 of 215,962 articles

Resilient Vision-Tabular Multimodal Learning under Modality Missingness

arXiv
Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete modality avai... read more 

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

arXiv
Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric deformations, because the model... read more 

Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection

arXiv
Zero-shot anomaly detection aims to identify defects in unseen categories without target-specific training. Existing methods usually apply the same feature transformation to all samples, treating normal and anomalous data uniformly despite their fund... read more 

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

arXiv
Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physical processes remains difficult, as models must combine object localiza... read more 

The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragments

arXiv
Jigsaw puzzle solving has been an increasingly popular task in the computer vision research community. Recent works have utilized cutting-edge architectures and computational approaches to reassemble groups of pieces into a coherent image, while achi... read more 

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

arXiv
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning:... read more 

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy

arXiv
RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning. In RL, diversity is often assumed to correlate with policy entropy, motivating entropy regularizat... read more 

Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning

arXiv
Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among prior approaches, sparse autoencoder (SAE)-based methods have attracted attention due to their abili... read more 

MULTI: Disentangling Camera Lens, Sensor, View, and Domain for Novel Image Generation

arXiv
Recent text-to-image models produce high-quality images, yet text ambiguity hinders precise control when specific styles or objects are required. There have been a number of recent works dealing with learning and composing multiple objects and patter... read more 

STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts

arXiv
We present STRUM (Spectral Transcription and Rhythm Understanding Model), an audio-to-chart pipeline that converts raw recordings into playable Clone Hero / YARG charts for drums, guitar, bass, vocals, and keys without any oracle metadata. STRUM is a... read more