Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 38,231 to 38,240 of 223,469 articles

CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation

arXiv
CLIP aligns image and text embeddings via contrastive learning and demonstrates strong zero-shot generalization. Its large-scale architecture requires substantial computational and memory resources, motivating the distillation of its capabilities int... read more 

CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation

arXiv
CLIP aligns image and text embeddings via contrastive learning and demonstrates strong zero-shot generalization. Its large-scale architecture requires substantial computational and memory resources, motivating the distillation of its capabilities int... read more 

Multimodal Dataset Distillation via Phased Teacher Models

arXiv
Multimodal dataset distillation aims to construct compact synthetic datasets that enable efficient compression and knowledge transfer from large-scale image-text data. However, existing approaches often fail to capture the complex, dynamically evolvi... read more 

PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders

arXiv
Vision Foundation Models (VFMs) pre-trained at scale enable a single frozen encoder to serve multiple downstream tasks simultaneously. Recent VFM-based encoder-only models for image and video segmentation, such as EoMT and VidEoMT, achieve competitiv... read more 

Shape and Substance: Dual-Layer Side-Channel Attacks on Local Vision-Language Models

arXiv
On-device Vision-Language Models (VLMs) promise data privacy via local execution. However, we show that the architectural shift toward Dynamic High-Resolution preprocessing (e.g., AnyRes) introduces an inherent algorithmic side-channel. Unlike static... read more 

Shape and Substance: Dual-Layer Side-Channel Attacks on Local Vision-Language Models

arXiv
On-device Vision-Language Models (VLMs) promise data privacy via local execution. However, we show that the architectural shift toward Dynamic High-Resolution preprocessing (e.g., AnyRes) introduces an inherent algorithmic side-channel. Unlike static... read more 

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

arXiv
Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we... read more 

CIAR: Interval-based Collaborative Decoding for Image Generation Acceleration

arXiv
Auto-regressive (AR) models have recently made notable progress in image generation, achieving performance comparable to diffusion-based approaches. However, their computational intensity and sequential nature impede on-device deployment, causing dis... read more 

Interpretable PM2.5 Forecasting for Urban Air Quality: A Comparative Study of Operational Time-Series Models

arXiv
Accurate short-term air-quality forecasting is essential for public health protection and urban management, yet many recent forecasting frameworks rely on complex, data-intensive, and computationally demanding models. This study investigates whether ... read more 

Knowledge-Guided Failure Prediction: Detecting When Object Detectors Miss Safety-Critical Objects

arXiv
Object detectors deployed in safety-critical environments can fail silently, e.g. missing pedestrians, workers, or other safety-critical objects without emitting any warning. Traditional Out Of Distribution (OOD) detection methods focus on identifyin... read more