Artificial Intelligence Medical Compendium

Explore the latest research on artificial intelligence and machine learning in medicine.

Showing 20,771 to 20,780 of 216,088 articles

MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation

arXiv
The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentrated on architectural optimizat... read more 

Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition

arXiv
Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The merging of these data domains represents a significant paradigm shift with far-reaching implications... read more 

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenizatio

arXiv
Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, existing methods universally extract features from only the last encoder layer, discard... read more 

Fixed-Point Neural Optimal Transport without Implicit Differentiation

arXiv
We propose an implicit neural formulation of optimal transport that eliminates adversarial min--max optimization and multi-network architectures commonly used in existing approaches. Our key idea is to parameterize a single potential in the Kantorovi... read more 

Muown: Row-Norm Control for Muon Optimization

arXiv
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon without decoupled weight decay, the spectral norm of weight matrices dri... read more 

Predicting 3D structure by latent posterior sampling

arXiv
The remarkable achievements of both generative models of 2D images and neural field representations for 3D scenes present a compelling opportunity to integrate the strengths of both approaches. In this work, we propose a methodology that combines a N... read more 

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

arXiv
Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which do not fully reflect continuous inspection processes in real industrial scenarios. We introduce MMV... read more 

Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories

arXiv
We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JEPA architectures have enabled latent-space planning in robotics and high-quality representation learning in vis... read more 

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

arXiv
Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a default safety layer for medical visual question answering (VQA). We argue that this practice is funda... read more 

Masked Generative Transformer Is What You Need for Image Editing

arXiv
Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications to propagate into areas that should remain intact. We propose a fundamentally different approach by... read more