Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 61,071 to 61,080 of 228,300 articles

RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation

arXiv
Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack ... read more 

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

arXiv
Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from them, or do not disclose their training data or recipe. As a result, ... read more 

Alterbute: Editing Intrinsic Attributes of Objects in Images

arXiv
We introduce Alterbute, a diffusion-based method for editing an object's intrinsic attributes in an image. We allow changing color, texture, material, and even the shape of an object, while preserving its perceived identity and scene context. Existin... read more 

DInf-Grid: A Neural Differential Equation Solver with Differentiable Feature Grids

arXiv
We present a novel differentiable grid-based representation for efficiently solving differential equations (DEs). Widely used architectures for neural solvers, such as sinusoidal neural networks, are coordinate-based MLPs that are both computationall... read more 

Future Optical Flow Prediction Improves Robot Control & Video Generation

arXiv
Future motion representations, such as optical flow, offer immense value for control and generative tasks. However, forecasting generalizable spatially dense motion representations remains a key challenge, and learning such forecasting from noisy, re... read more 

ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research

arXiv
Current progress in out-of-distribution (OOD) detection is limited by the lack of large, high-quality datasets with clearly defined OOD categories across varying difficulty levels (near- to far-OOD) that support both fine- and coarse-grained computer... read more 

SurfSLAM: Sim-to-Real Underwater Stereo Reconstruction For Real-Time SLAM

arXiv
Localization and mapping are core perceptual capabilities for underwater robots. Stereo cameras provide a low-cost means of directly estimating metric depth to support these tasks. However, despite recent advances in stereo depth estimation on land, ... read more 

Can Vision-Language Models Understand Construction Workers? An Exploratory Study

arXiv
As robotics become increasingly integrated into construction workflows, their ability to interpret and respond to human behavior will be essential for enabling safe and effective collaboration. Vision-Language Models (VLMs) have emerged as a promisin... read more 

SecMLOps: A Comprehensive Framework for Integrating Security Throughout the MLOps Lifecycle

arXiv
Machine Learning (ML) has emerged as a pivotal technology in the operation of large and complex systems, driving advancements in fields such as autonomous vehicles, healthcare diagnostics, and financial fraud detection. Despite its benefits, the depl... read more 

AI-Guided Human-In-the-Loop Inverse Design of High Performance Engineering Structures

arXiv
Inverse design tools such as Topology Optimization (TO) can achieve new levels of improvement for high-performance engineered structures. However, widespread use is hindered by high computational times and a black-box nature that inhibits user intera... read more