Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 60,321 to 60,330 of 228,072 articles

Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers

arXiv
Vision-Language Models (VLMs) are rapidly replacing unimodal encoders in modern retrieval and recommendation systems. While their capabilities are well-documented, their robustness against adversarial manipulation in competitive ranking scenarios rem... read more 

CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training

arXiv
The functions of different regions of the human brain are closely linked to their distinct cytoarchitecture, which is defined by the spatial arrangement and morphology of the cells. Identifying brain regions by their cytoarchitecture enables various ... read more 

SDiT: Semantic Region-Adaptive for Diffusion Transformers

arXiv
Diffusion Transformers (DiTs) achieve state-of-the-art performance in text-to-image synthesis but remain computationally expensive due to the iterative nature of denoising and the quadratic cost of global attention. In this work, we observe that deno... read more 

OpenNavMap: Structure-Free Topometric Mapping via Large-Scale Collaborative Localization

arXiv
Scalable and maintainable map representations are fundamental to enabling large-scale visual navigation and facilitating the deployment of robots in real-world environments. While collaborative localization across multi-session mapping enhances effic... read more 

Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations

arXiv
Deep learning has achieved remarkable success in image recognition, yet their inherent opacity poses challenges for deployment in critical domains. Concept-based interpretations aim to address this by explaining model reasoning through human-understa... read more 

A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models

arXiv
Vision-language pre-training (VLP) models are vulnerable to adversarial examples, particularly in black-box scenarios. Existing multimodal attacks often suffer from limited perturbation diversity and unstable multi-stage pipelines. To address these c... read more 

Adaptive Multi-Scale Correlation Meta-Network for Few-Shot Remote Sensing Image Classification

arXiv
Few-shot learning in remote sensing remains challenging due to three factors: the scarcity of labeled data, substantial domain shifts, and the multi-scale nature of geospatial objects. To address these issues, we introduce Adaptive Multi-Scale Correl... read more 

CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding

arXiv
Surgical action triplet recognition aims to understand fine-grained surgical behaviors by modeling the interactions among instruments, actions, and anatomical targets. Despite its clinical importance for workflow analysis and skill assessment, progre... read more 

EmoKGEdit: Training-free Affective Injection via Visual Cue Transformation

arXiv
Existing image emotion editing methods struggle to disentangle emotional cues from latent content representations, often yielding weak emotional expression and distorted visual structures. To bridge this gap, we propose EmoKGEdit, a novel training-fr... read more 

FlowIID: Single-Step Intrinsic Image Decomposition via Latent Flow Matching

arXiv
Intrinsic Image Decomposition (IID) separates an image into albedo and shading components. It is a core step in many real-world applications, such as relighting and material editing. Existing IID models achieve good results, but often use a large num... read more