Latest AI and machine learning research in medical education for healthcare professionals.
Compositional scene reconstruction seeks to create object-centric representations rather than holistic scenes from real-world videos, which is natively applicable for simulation and interaction. Conventional compositional reconstruction approaches primarily emphasize on visual appearance and show limited generalization ability to real-world scenarios. In this paper, we propose SimRecon, a framewor...
Instruction-based video editing has witnessed rapid progress, yet current methods often struggle with precise visual control, as natural language is inherently limited in describing complex visual nuances. Although reference-guided editing offers a robust solution, its potential is currently bottlenecked by the scarcity of high-quality paired training data. To bridge this gap, we introduce a scala...
Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a p...
Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some requir...
Background Clinicians in care management programs are often in low supply relative to patient demand, especially in US Medicaid programs, and must sim...
Background: Objective Structured Clinical Examination (OSCE; Clinical Performance Examination [CPX] in South Korea) is a high-stakes assessment of cli...
360 panoramic images are increasingly used in virtual reality, autonomous driving, and robotics for holistic scene understanding. However, current Vis...
Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some requir...
The inability of Large Language Models (LLMs) to modulate their personality expression in response to evolving dialogue dynamics hinders their perform...
Uniform-state discrete diffusion models excel at few-step generation and guidance due to their ability to self-correct, making them preferred over aut...
The growing prevalence of tampered images poses serious security threats, highlighting the urgent need for reliable detection methods. Multimodal larg...
Recent advances in garment pattern generation have shown promising progress. However, existing feed-forward methods struggle with diverse poses and vi...
Developing foundation models for electroencephalography (EEG) remains challenging due to the signal's low signal-to-noise ratio and complex spectro-te...
Electrocardiograms (ECG) are electrical recordings of the heart that are critical for diagnosing cardiovascular conditions. ECG language models (ELMs)...
We address fine-grained visual reasoning in multimodal large language models (MLLMs), where key evidence may reside in tiny objects, cluttered regions...
3D Gaussian Splatting (3DGS) has emerged as a powerful approach for novel view synthesis. However, the number of Gaussian primitives often grows subst...
Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF)....
Introduction Eighteenth century medical texts document a formative period in the evolution of clinical reasoning, yet their integration into modern me...
Procedural generation techniques in 3D rendering engines have revolutionized the creation of complex environments, reducing reliance on manual design....
Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss o...