Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 43,711 to 43,720 of 224,055 articles

VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer

arXiv
Zero-shot anomaly detection (ZSAD) requires detecting and localizing anomalies without access to target-class anomaly samples. Mainstream methods rely on vision-language models (VLMs) such as CLIP: they build hand-crafted or learned prompt sets for n... read more 

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

arXiv
As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands... read more 

MJ1: Multimodal Judgment via Grounded Verification

arXiv
Multimodal judges struggle to ground decisions in visual evidence. We present MJ1, a multimodal judge trained with reinforcement learning that enforces visual grounding through a structured grounded verification chain (observations $\rightarrow$ clai... read more 

ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation

arXiv
Existing aerial Vision-Language Navigation (VLN) methods predominantly adopt a detection-and-planning pipeline, which converts open-vocabulary detections into discrete textual scene graphs. These approaches are plagued by inadequate spatial reasoning... read more 

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

arXiv
Infrared-visible (IR-VIS) image fusion is vital for perception and security, yet most methods rely on the availability of both modalities during training and inference. When the infrared modality is absent, pixel-space generative substitutes become h... read more 

VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion

arXiv
Generating realistic cast shadows for inserted foreground objects is a crucial yet challenging problem in image composition, where maintaining geometric consistency of shadow and object in complex scenes remains difficult due to the ill-posed nature ... read more 

Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades

arXiv
Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit pose-based c... read more 

QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration

arXiv
Real-world image restoration (RWIR) is a highly challenging task due to the absence of clean ground-truth images. Many recent methods resort to pseudo-label (PL) supervision, often within a Mean-Teacher (MT) framework. However, these methods face a c... read more 

Speed3R: Sparse Feed-forward 3D Reconstruction Models

arXiv
While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a prohibitive computatio... read more 

See and Switch: Vision-Based Branching for Interactive Robot-Skill Programming

arXiv
Programming robots by demonstration (PbD) is an intuitive concept, but scaling it to real-world variability remains a challenge for most current teaching frameworks. Conditional task graphs are very expressive and can be defined incrementally, which ... read more