Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 37,391 to 37,400 of 223,469 articles

On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers

arXiv
Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt. This typicality bias presents a ch... read more 

PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models

arXiv
Acquiring labeled datasets for 3D human mesh estimation is challenging due to depth ambiguities and the inherent difficulty of annotating 3D geometry from monocular images. Existing datasets are either real, with manually annotated 3D geometry and li... read more 

Gen-Searcher: Reinforcing Agentic Search for Image Generation

arXiv
Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally constrained by frozen internal knowledge, thus often failing on real-world scenarios that are knowled... read more 

Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey

arXiv
Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While powerful, this increased expr... read more 

ForestSim: A Synthetic Benchmark for Intelligent Vehicle Perception in Unstructured Forest Environments

arXiv
Robust scene understanding is essential for intelligent vehicles operating in natural, unstructured environments. While semantic segmentation datasets for structured urban driving are abundant, the datasets for extremely unstructured wild environment... read more 

JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

arXiv
Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese is included in several multil... read more 

MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation

arXiv
Modern generative models have demonstrated the ability to solve challenging mathematical problems. In many real-world settings, however, mathematical solutions must be expressed visually through diagrams, plots, geometric constructions, and structure... read more 

Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs

arXiv
Image-to-point-cloud (I2P) registration aims to align 2D images with 3D point clouds by establishing reliable 2D-3D correspondences. The drastic modality gap between images and point clouds makes it challenging to learn features that are both discrim... read more 

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

arXiv
Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending this lear... read more 

RetinexDualV2: Physically-Grounded Dual Retinex for Generalized UHD Image Restoration

arXiv
We propose RetinexDualV2, a unified, physically grounded dual-branch framework for diverse Ultra-High-Definition (UHD) image restoration. Unlike generic models, our method employs a Task-Specific Physical Grounding Module (TS-PGM) to extract degradat... read more