Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 41,571 to 41,580 of 223,853 articles

CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models

arXiv
We present CyCLeGen, a unified vision-language foundation model capable of both image understanding and image generation within a single autoregressive framework. Unlike existing vision models that depend on separate modules for perception and synthe... read more 

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

arXiv
Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometr... read more 

Thermal Image Refinement with Depth Estimation using Recurrent Networks for Monocular ORB-SLAM3

arXiv
Autonomous navigation in GPS-denied and visually degraded environments remains challenging for unmanned aerial vehicles (UAVs). To this end, we investigate the use of a monocular thermal camera as a standalone sensor on a UAV platform for real-time d... read more 

Edit2Interp: Adapting Image Foundation Models from Spatial Editing to Video Frame Interpolation with Few-Shot Learning

arXiv
Pre-trained image editing models exhibit strong spatial reasoning and object-aware transformation capabilities acquired from billions of image-text pairs, yet they possess no explicit temporal modeling. This paper demonstrates that these spatial prio... read more 

Empowering Chemical Structures with Biological Insights for Scalable Phenotypic Virtual Screening

arXiv
Motivation: The scalable identification of bioactive compounds is essential for contemporary drug discovery. This process faces a key trade-off: structural screening offers scalability but lacks biological context, whereas high-content phenotypic pro... read more 

MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal

arXiv
Memes represent a tightly coupled, multimodal form of social expression, in which visual context and overlaid text jointly convey nuanced affect and commentary. Inspired by cognitive reappraisal in psychology, we introduce Meme Reappraisal, a novel m... read more 

One CT Unified Model Training Framework to Rule All Scanning Protocols

arXiv
Non-ideal measurement computed tomography (NICT), which lowers radiation at the cost of image quality, is expanding the clinical use of CT. Although unified models have shown promise in NICT enhancement, most methods require paired data, which is an ... read more 

Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

arXiv
Following major advances in text and image generation, the video domain has surged, producing highly realistic and controllable sequences. Along with this progress, these models also raise serious concerns about misinformation, making reliable detect... read more 

Rethinking Machine Unlearning: Models Designed to Forget via Key Deletion

arXiv
Machine unlearning is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training samples. Despite this, most existing methods tackle the problem purely from a post-hoc pe... read more 

GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents

arXiv
Recent progress in Multimodal Large Language Models (MLLMs) has enabled mobile GUI agents capable of visual perception, cross-modal reasoning, and interactive control. However, existing benchmarks are largely English-centric and fail to capture the l... read more