Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 17,161 to 17,170 of 213,726 articles

Improved Baselines with Representation Autoencoders

arXiv
Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insights which simplify and improve RAE. First, we study a generalized formu... read more 

CineMatte: Background Matting for Virtual Production and Beyond

arXiv
LED Virtual Production (VP) uses large LED volumes to render backgrounds in real time, enabling in-camera visual effects but making post-shot changes labor-intensive. We address this with CineMatte, a robust background matting framework for VP and be... read more 

Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation

arXiv
Ensemble disagreement is widely used as a proxy for epistemic uncertainty in medical image segmentation. In practice, many studies form ensembles via K-fold cross-validation (CV), yet refer to them as ``deep ensembles'' (DE). Because CV members are t... read more 

Vision Foundation Models as Generalist Tokenizers for Image Generation

arXiv
In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenizer, we utilize a frozen VFM as the encoder and introduce two key innova... read more 

NEWTON: Agentic Planning for Physically Grounded Video Generation

arXiv
Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: text prompts are lossy compressio... read more 

Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

arXiv
Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed to accurately assess structural integrity, but progress in defect segmentation for civil infrastru... read more 

Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology

arXiv
Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology. However, fine-tuning billions of parameters on scarce, expert-annotated pathology data is prohibit... read more 

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval

arXiv
E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modality disparity -- a visual query... read more 

What is Holding Back Latent Visual Reasoning?

arXiv
Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired by this, several works on Vision-Language Models have recently explored chain-of-thought reasoning wi... read more 

NeRF-based Spacecraft Reconstruction from Close-Range Monocular Imagery Under Illumination Variability and Pose Uncertainty

arXiv
Autonomous rendezvous and proximity operations around uncooperative, unknown spacecraft are critical for active debris removal and on-orbit servicing missions. A key component of such operations is the offline reconstruction of a 3D model of the targ... read more