Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 54,871 to 54,880 of 226,183 articles

Deep Multivariate Models with Parametric Conditionals

arXiv
We consider deep multivariate models for heterogeneous collections of random variables. In the context of computer vision, such collections may e.g. consist of images, segmentations, image attributes, and latent variables. When developing such models... read more 

Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images

arXiv
Open-vocabulary object detection in remote sensing commonly relies on text-only prompting to specify target categories, implicitly assuming that inference-time category queries can be reliably grounded through pretraining-induced text-visual alignmen... read more 

Grounding Generated Videos in Feasible Plans via World Models

arXiv
Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to executable actions. To addres... read more 

Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If Calibrated

arXiv
Despite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake sam... read more 

Enhancing Multi-Image Understanding through; Scaling

arXiv
Large Vision-Language Models (LVLMs) achieve strong performance on single-image tasks, but their performance declines when multiple images are provided as input. One major reason is the cross-image information leakage, where the model struggles to di... read more 

Leveraging Latent Vector Prediction for Localized Control in Image Generation via Diffusion Models

arXiv
Diffusion models emerged as a leading approach in text-to-image generation, producing high-quality images from textual descriptions. However, attempting to achieve detailed control to get a desired image solely through text remains a laborious trial-... read more 

SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors

arXiv
Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3D scenes using 3D Gaussian Spl... read more 

SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors

arXiv
Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3D scenes using 3D Gaussian Spl... read more 

UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving

arXiv
World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera video or LiDAR sequence synthesi... read more 

ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning

arXiv
Large multimodal reasoning models solve challenging visual problems via explicit long-chain inference: they gather visual clues from images and decode clues into textual tokens. Yet this capability also increases hallucinations, where the model gener... read more