Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 44,351 to 44,360 of 224,055 articles

Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models

arXiv
Recent advances in large-scale pretrained vision models have demonstrated impressive capabilities across a wide range of downstream tasks, including cross-modal and multi-modal scenarios. However, their direct application to 3D human skeleton data re... read more 

Imagine How To Change: Explicit Procedure Modeling for Change Captioning

arXiv
Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring the rich temporal dynamics of the change procedure, which is the key ... read more 

Towards High-resolution and Disentangled Reference-based Sketch Colorization

arXiv
Sketch colorization is a critical task for automating and assisting in the creation of animations and digital illustrations. Previous research identified the primary difficulty as the distribution shift between semantically aligned training data and ... read more 

Technical Report: Automated Optical Inspection of Surgical Instruments

arXiv
In the dynamic landscape of modern healthcare, maintaining the highest standards in surgical instruments is critical for clinical success. This report explores the diverse realm of surgical instruments and their associated manufacturing defects, emph... read more 

MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs

arXiv
Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across variables. Existing ISTS forecasting methods often solely utilize historical observations to predict f... read more 

RePer-360: Releasing Perspective Priors for 360$^\circ$ Depth Estimation via Self-Modulation

arXiv
Recent depth foundation models trained on perspective imagery achieve strong performance, yet generalize poorly to 360$^\circ$ images due to the substantial geometric discrepancy between perspective and panoramic domains. Moreover, fully fine-tuning ... read more 

Demystifying KAN for Vision Tasks: The RepKAN Approach

arXiv
Remote sensing image classification is essential for Earth observation, yet standard CNNs and Transformers often function as uninterpretable black-boxes. We propose RepKAN, a novel architecture that integrates the structural efficiency of CNNs with t... read more 

ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning

arXiv
Multi-view spatial reasoning remains difficult for current vision-language models. Even when multiple viewpoints are available, models often underutilize cross-view relations and instead rely on single-image shortcuts, leading to fragile performance ... read more 

StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision

arXiv
Reasoning-based text-to-image (T2I) generation requires models to interpret complex prompts accurately. Existing reasoning frameworks can be broadly categorized into two types: (1) Text-Only Reasoning, which is computationally efficient but lacks acc... read more 

Ensemble Learning with Sparse Hypercolumns

arXiv
Directly inspired by findings in biological vision, high-dimensional hypercolumns are feature vectors built by concatenating multi-scale activations of convolutional neural networks for a single image pixel location. Together with powerful classifier... read more