Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 39,381 to 39,390 of 223,737 articles

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

arXiv
Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface fluid migration. Understanding these phenomena is crucial for assessing h... read more 

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

arXiv
Multi-Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW-13 and RefCOCO. However, state-of-the-art models still struggle to generalize to out-of-distribution classes, tasks and imagin... read more 

InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting

arXiv
Recent diffusion-based models achieve photorealism in image inpainting but require many sampling steps, limiting practical use. Few-step text-to-image models offer faster generation, but naively applying them to inpainting yields poor harmonization a... read more 

One View Is Enough! Monocular Training for In-the-Wild Novel View Generation

arXiv
Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is not necessary: one view is enough. We present OVIE, trained entirely on unpaired internet images. We l... read more 

Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation

arXiv
Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming video generation. The growing demand for higher resolutions, frame rates, and context lengths, however,... read more 

VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

arXiv
Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on cha... read more 

DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models

arXiv
Optical flow models trained on high-quality data often degrade severely when confronted with real-world corruptions such as blur, noise, and compression artifacts. To overcome this limitation, we formulate Degradation-Aware Optical Flow, a new task t... read more 

UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation

arXiv
Unified models capable of interleaved generation have emerged as a promising paradigm, with the community increasingly converging on autoregressive modeling for text and flow matching for image generation. To advance this direction, we propose a unif... read more 

MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage

arXiv
Vision Language Models (VLMs) are increasingly used for tasks like medical report generation and visual question answering. However, fluent diagnostic text does not guarantee safe visual understanding. In clinical practice, interpretation begins with... read more 

OccAny: Generalized Unconstrained Urban 3D Occupancy

arXiv
Relying on in-domain annotations and precise sensor-rig priors, existing 3D occupancy prediction methods are limited in both scalability and out-of-domain generalization. While recent visual geometry foundation models exhibit strong generalization ca... read more