Public Health & Policy

Medical Education

Latest AI and machine learning research in medical education for healthcare professionals.

4,138 articles
Stay Ahead - Weekly Medical Education research updates
Subscribe
Browse Categories
Showing 1201-1220 of 4,138 articles

On the Design Fundamentals of Pixel Text Representation Learning

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual text understanding. In this work, we investigate the fundamental design principles required for robust visual text representation learning....

Sep 1 2026 2609.01147v1

Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison

Self-supervised pre-training via cross-view completion learns strong features for 3D vision from co-visible regions of image pairs. However, the reference view provides little information for reconstructing non-co-visible patches, implicitly yielding a monocular training signal in these regions. We introduce Gekko, which turns this limitation into a useful signal. The relative improvement of the c...

Sep 1 2026 2609.01530v1
SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling

Histopathological whole slide images (WSIs) are central to cancer diagnosis, but their gigapixel scale, tissue heterogeneity, weak slide-level supervi...

Aug 31 2026 2609.00396v1
Prospective In-silico Simulation of the VESALIUS-CV Trial Using Biomedical Knowledge Graph and Real-World Data-Driven AI Modeling

Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulati...

Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment

Federated learning (FL) has emerged as a key approach for training models across decentralized data, yet benchmarking in FL remains difficult to repro...

Aug 25 2026 2608.25114v1
NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation

The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the real world. However, transforming raw visual obs...

Aug 25 2026 2608.24212v1
Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness

Imputing physiological time series (arterial blood pressure, blood glucose, etc.) is essential for addressing the missingness that pervades clinical d...

Aug 21 2026 2608.21207v1
Exploring the Performance Frontier of Compact Unified Image Generation Models

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore ho...

Aug 20 2026 2608.20334v2
Far from the Crowd: Scalable Self-Supervised Learning via Geographic Isolation

Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and...

Aug 20 2026 2608.19766v1
CalcSeg: Confidence-aware 3D Latent Context Curriculum Learning For Myocardial Scar Segmentation From Single-Stack LGE-CMRs

Myocardial scar segmentation from single-stack late gadolinium-enhanced cardiac magnetic resonance (LGE-CMR) imaging has been a longstanding and clini...

Aug 20 2026 2608.20305v1
Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore ho...

Aug 20 2026 2608.20334v1
GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual e...

Aug 19 2026 2608.19355v1
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically o...

Aug 18 2026 2608.18076v1
Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors

Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with ...

Aug 13 2026 2608.12746v1
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent ...

Aug 11 2026 2608.10522v1
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates signif...

Aug 11 2026 2608.10954v1
SAR2Agri: Learning SAR Intensity Representations for Agricultural Monitoring

Agricultural monitoring faces unique challenges, arising from the landscape's complex temporal, phenological, and climate dynamics, yet monitoring the...

Aug 11 2026 2608.11142v1
Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware

IoT firmware vulnerability detection remains challenging due to heterogeneous firmware ecosystems, resource-constrained platforms, and limitations in ...

Aug 11 2026 2608.11492v1
Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f...

Aug 9 2026 2608.08545v1
Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Mode...

Aug 9 2026 2608.08964v1
Browse Categories