Latest AI and machine learning research in alternative medicine for healthcare professionals.
Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck by scoring token importance and pruning low-scoring tokens in a single pass. However, one-shot scoring is insufficient because a token's prompt-relevant usefulness depen...
Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textualizes that evidence and then applies a text predictor. We ask: how should visual ...
The emergence of medical deepfakes, i.e., medical images manipulated by deep generative models, poses a significant threat to clinical workflows. Howe...
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both enti...
Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions ...
Problem Authentic patient encounters are the raw material of clinical learning, yet the educational resources learners receive are rarely keyed to the...
Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insuffici...
Rapid advances in image generation models call for interpretable AI-generated image detection methods that not only determine authenticity but also pr...
Tracking links observations of the same object through visual change, yet cannot by itself determine when the object is empty or filled, intact or cut...
Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens i...
The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked exp...
Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. E...
Decoding visual experience from non-invasive brain activity is central to neuroscience and brain-computer interfaces. Functional magnetic resonance im...
Interaction between users and LLM agents is increasingly multimodal: conversations interleave text with images, and a later question may target either...
Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimoda...
AI-generated image forgeries are becoming increasingly realistic and difficult to characterize with fixed manipulation patterns. As generative models ...
In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypi...
A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study ...
Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-oc...
Here is the plain text version optimized for arXiv's submission form. Custom macros (like \CV and \SI) have been converted to standard text/math so th...