Psychiatry

Schizophrenia

Latest AI and machine learning research in schizophrenia for healthcare professionals.

3,224 articles
Stay Ahead - Weekly Schizophrenia research updates
Subscribe
Browse Categories
Showing 701-720 of 3,224 articles

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs

Multimodal large language models (MLLMs) have become a key interface for visual reasoning and grounded question answering, yet they remain vulnerable to visual hallucinations, where generated responses contradict image content or mention nonexistent objects. A central challenge is that hallucination is not always caused by a simple lack of visual attention: the model may still assign substantial a...

May 12 2026 2605.11559v1

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before inference. We study whether grounding improves when the visual interface is moved closer to the underlying camera measurement. We formulate measurement-grounded vision-language learning and instantiate it as PRISM-VL, which combines RAW-derived Meas.-XYZ i...

May 12 2026 2605.11727v1
CAWI: Copula-Aligned Weight Initialization for Randomized Neural Networks

Randomized neural networks (RdNNs) enable efficient, backpropagation-free training by freezing randomly initialized input-to-hidden weights, which per...

May 12 2026 2605.12580v1
CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis

Foundation diffusion models can generate photorealistic natural images, but adapting them to medical imaging remains challenging. In medical adaptatio...

May 12 2026 2605.12650v1
Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet in...

May 11 2026 2605.10002v1
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs

Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual t...

May 11 2026 2605.10050v1
Claim-Level Transparency Analysis of LLM-Generated Diagnostic Reports: A Metabolic and Endocrine Biomarker Study

Large language models are increasingly deployed in clinical decision-support contexts, yet systematic evaluation of their factual reliability in gener...

Fast and Ultra-Capable Protein Design: Advancing the Frontier Through Atomistic SE(3)-Equivariance with Genie 3

Despite the breakneck pace of progress in protein design methodology, frontier problems remain challenging, with leading methods struggling to design ...

MK-ResRecon: Multi-Kernel Residual Framework for Texture-Aware 3D MRI Refinement from Sparse 2D Slices

Magnetic Resonance Imaging (MRI) acquisition remains a time-intensive and patient-straining process, as prolonged scan dura- tions increase the likeli...

May 5 2026 2605.03432v1
FluxFlow: Conservative Flow-Matching for Astronomical Image Super-Resolution

Ground-to-space astronomical super-resolution requires recovering space-quality images from ground-based observations that are simultaneously limited ...

May 5 2026 2605.03749v1
Online Self-Calibration Against Hallucination in Vision-Language Models

Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image...

May 1 2026 2605.00323v1
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning

Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently genera...

Apr 28 2026 2604.25809v1
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters

Vision-language models (VLMs) are increasingly deployed as evaluators in tasks requiring nuanced image understanding, yet their reliability in scoring...

Apr 27 2026 2604.24346v1
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding

Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding,...

Apr 24 2026 2604.22498v1
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry

Vision evaluations are typically done through multi-step processes. In most contemporary fields, experts analyze images using structured, evidence-bas...

Apr 22 2026 2604.20983v1
Hallucination Early Detection in Diffusion Models

Text-to-Image generation has seen significant advancements in output realism with the advent of diffusion models. However, diffusion models encounter ...

Apr 22 2026 2604.20354v1
Improving clinical interpretability of linear neuroimaging models through feature whitening

Linear models are widely used in computational neuroimaging to identify biomarkers associated with brain pathologies. However, interpreting the learne...

Apr 22 2026 2604.20675v1
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they s...

Apr 22 2026 2604.20696v1
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps

Hallucinations in Speech Large Language Models (SpeechLLMs) pose significant risks, yet existing detection methods typically rely on gold-standard out...

Apr 21 2026 2604.19565v1
Lucky High Dynamic Range Smartphone Imaging

While the human eye can perceive an impressive twenty stops of dynamic range, smartphone camera sensors remain limited to about twelve stops despite d...

Apr 21 2026 2604.19976v1
Browse Categories