Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 41,311 to 41,320 of 223,853 articles

What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers

arXiv
Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as positional encoding) can lead to these models displaying positional b... read more 

Dynamic Meta-Layer Aggregation for Byzantine-Robust Federated Learning

arXiv
Federated Learning (FL) is increasingly applied in sectors like healthcare, finance, and IoT, enabling collaborative model training while safeguarding user privacy. However, FL systems are susceptible to Byzantine adversaries that inject malicious up... read more 

SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation

arXiv
Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpected artifacts, but instead can... read more 

ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K

arXiv
Learning in simulation provides a useful foundation for scaling robotic manipulation capabilities. However, this paradigm often suffers from a lack of data-generation-ready digital assets, in both scale and diversity. In this work, we present ManiTwi... read more 

MessyKitchens: Contact-rich object-level 3D scene reconstruction

arXiv
Monocular 3D scene reconstruction has recently seen significant progress. Powered by the modern neural architectures and large-scale data, recent methods achieve high performance in depth estimation from a single image. Meanwhile, reconstructing and ... read more 

Non-perturbative Bacterial Identification Directly from Solid Agar Plates Using Raman

arXiv
Raman spectroscopy is a promising tool for microbial identification, yet its implementation in microbiology and clinical workflow is still restricted due to the accompanying additional preparation required to focus on microbial signals. Here, we demo... read more 

PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models

arXiv
Vision-Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required for manipulation remains limited. In particular, estimating the mass of real-world objects is essen... read more 

Topology-Guided Biomechanical Profiling: A White-Box Framework for Opportunistic Screening of Spinal Instability on Routine CT

arXiv
Routine oncologic computed tomography (CT) presents an ideal opportunity for screening spinal instability, yet prophylactic stabilization windows are frequently missed due to the complex geometric reasoning required by the Spinal Instability Neoplast... read more 

CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization

arXiv
Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, ... read more 

MSRAMIE: Multimodal Structured Reasoning Agent for Multi-instruction Image Editing

arXiv
Existing instruction-based image editing models perform well with simple, single-step instructions but degrade in realistic scenarios that involve multiple, lengthy, and interdependent directives. A main cause is the scarcity of training data with co... read more