Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 46,231 to 46,240 of 224,055 articles

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

arXiv
Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their feature representations are poorly aligned across different modalities. For instance, the feature embedding for an RGB... read more 

GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction

arXiv
Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video generation for avatar reconstruct... read more 

Manifold-Preserving Superpixel Hierarchies and Embeddings for the Exploration of High-Dimensional Images

arXiv
High-dimensional images, or images with a high-dimensional attribute vector per pixel, are commonly explored with coordinated views of a low-dimensional embedding of the attribute space and a conventional image representation. Nowadays, such images c... read more 

RAViT: Resolution-Adaptive Vision Transformer

arXiv
Vision transformers have recently made a breakthrough in computer vision showing excellent performance in terms of precision for numerous applications. However, their computational cost is very high compared to alternative approaches such as Convolut... read more 

What You Read is What You Classify: Highlighting Attributions to Text and Text-Like Inputs

arXiv
At present, there are no easily understood explainable artificial intelligence (AI) methods for discrete token inputs, like text. Most explainable AI techniques do not extend well to token sequences, where both local and global features matter, becau... read more 

HumanOrbit: 3D Human Reconstruction as 360° Orbit Generation

arXiv
We present a method for generating a full 360° orbit video around a person from a single input image. Existing methods typically adapt image-based diffusion models for multi-view synthesis, but yield inconsistent results across views and with the ori... read more 

Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

arXiv
Decoupled dataset distillation (DD) compresses large corpora into a few synthetic images by matching a frozen teacher's statistics. However, current residual-matching pipelines rely on static real patches, creating a fit-complexity gap and a pull-to-... read more 

AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation

arXiv
The expansion of retrieval-augmented generation (RAG) into multimodal domains has intensified the challenge for processing complex visual documents, such as financial reports. While page-level chunking and retrieval is a natural starting point, it cr... read more 

Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

arXiv
Vision-language models (VLMs) show promise in drafting radiology reports, yet they frequently suffer from logical inconsistencies, generating diagnostic impressions unsupported by their own perceptual findings or missing logically entailed conclusion... read more 

DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

arXiv
Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in... read more