Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 38,201 to 38,210 of 223,469 articles

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

arXiv
Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Net and dual-stream architectures (e.g., Flux), this task remains under-explored in the recent emerging ... read more 

Learning domain-invariant features through channel-level sparsification for Out-Of Distribution Generalization

arXiv
Out-of-Distribution (OOD) generalization has become a primary metric for evaluating image analysis systems. Since deep learning models tend to capture domain-specific context, they often develop shortcut dependencies on these non-causal features, lea... read more 

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

arXiv
We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces three maj... read more 

Pixelis: Reasoning in Pixels, from Seeing to Acting

arXiv
Most vision-language systems are static observers: they describe pixels, do not act, and cannot safely improve under shift. This passivity limits generalizable, physically grounded visual intelligence. Learning through action, not static description,... read more 

Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning

arXiv
Multimodal learning integrates complementary information from different modalities such as image, text, and audio to improve model performance, but its success relies on large-scale labeled data, which is costly to obtain. Active learning (AL) mitiga... read more 

MoireMix: A Formula-Based Data Augmentation for Improving Image Classification Robustness

arXiv
Data augmentation is a key technique for improving the robustness of image classification models. However, many recent approaches rely on diffusion-based synthesis or complex feature mixing strategies, which introduce substantial computational overhe... read more 

EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions

arXiv
Smart glass is emerging as an useful device since it provides plenty of insights under hands-busy, eyes-on-task situations. To understand the context of the wearer, 6D object pose estimation in egocentric view is becoming essential. However, existing... read more 

Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models

arXiv
Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fixed-length token compression, disrupting volume... read more 

Vision Hopfield Memory Networks

arXiv
Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress, enabling unified modeling across images, text, and beyond. Despite their empirical success, these ar... read more 

Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling

arXiv
In complex environments, infrared object detection exhibits broad applicability and stability across diverse scenarios. However, infrared object detection is vulnerable to both common corruptions and adversarial examples, leading to potential securit... read more