Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 31,451 to 31,460 of 220,544 articles

Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation

arXiv
Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge arises from modality heterogeneity: many clinical sites possess only a subset of modalities due to ... read more 

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts

arXiv
Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cluster of high- and mid-resource scripts. We introduce GlotOCR Bench, a comprehensive benchmark eva... read more 

See, Point, Refine: Multi-Turn Approach to GUI Grounding with Visual Feedback

arXiv
Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen actions, but editing-level grounding in dense coding interfaces, where sub-pixel accuracy is required ... read more 

Representation geometry shapes task performance in vision-language modeling for CT enterography

arXiv
Computed tomography (CT) enterography is a primary imaging modality for assessing inflammatory bowel disease (IBD), yet the representational choices that best support automated analysis of this modality are unknown. We present the first study of visi... read more 

Conflated Inverse Modeling to Generate Diverse and Temperature-Change Inducing Urban Vegetation Patterns

arXiv
Urban areas are increasingly vulnerable to thermal extremes driven by rapid urbanization and climate change. Traditionally, thermal extremes have been monitored using Earth-observing satellites and numerical modeling frameworks. For example, land sur... read more 

Visual Preference Optimization with Rubric Rewards

arXiv
The effectiveness of Direct Preference Optimization (DPO) depends on preference data that reflect the quality differences that matter in multimodal tasks. Existing pipelines often rely on off-policy perturbations or coarse outcome-based signals, whic... read more 

Generative Refinement Networks for Visual Synthesis

arXiv
While diffusion models dominate the field of visual generation, they are computationally inefficient, applying a uniform computational effort regardless of different complexity. In contrast, autoregressive (AR) models are inherently complexity-aware,... read more 

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis

arXiv
Large Language Models (LLMs) and Vision-Language Models (VLMs) increasingly generate indoor scenes through intermediate structures such as layouts and scene graphs, yet evaluation still relies on LLM or VLM judges that score rendered views, making ju... read more 

Multi-modal panoramic 3D outdoor datasets for place categorization

arXiv
We present two multi-modal panoramic 3D outdoor (MPO) datasets for semantic place categorization with six categories: forest, coast, residential area, urban area and indoor/outdoor parking lot. The first dataset consists of 650 static panoramic scans... read more 

PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction

arXiv
3D Gaussian Splatting (3DGS) has recently enabled highly photorealistic 3D reconstruction from casually captured multi-view images. However, this accessibility raises a privacy concern: publicly available images or videos can be exploited to reconstr... read more