Psychiatry

Schizophrenia

Latest AI and machine learning research in schizophrenia for healthcare professionals.

3,500 articles
Stay Ahead - Weekly Schizophrenia research updates
Subscribe
Browse Categories
Showing 1101-1120 of 3,500 articles

PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is designed to evaluate both the accuracy a...

CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model

Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextualization (CC$^2$) problem through thr...

Treble Counterfactual VLMs: A Causal Approach to Hallucination

Vision-Language Models (VLMs) have advanced multi-modal tasks like image captioning, visual question answering, and reasoning. However, they often g...

SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner

Large Visual Language Models (LVLMs) increasingly rely on preference alignment to ensure reliability, which steers the model behavior via preference...

TPC: Cross-Temporal Prediction Connection for Vision-Language Model Hallucination Reduction

Vision-language models (VLMs) have achieved remarkable advancements, capitalizing on the impressive capabilities of large language models (LLMs) acr...

Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias

Score-based diffusion models have achieved incredible performance in generating realistic images, audio, and video data. While these models produce ...

See What You Are Told: Visual Attention Sink in Large Multimodal Models

Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally...

WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation

Object Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although rec...

MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models

Large Vision Language Models (LVLMs) are becoming increasingly important in the medical domain, yet Medical LVLMs (Med-LVLMs) frequently generate ha...

The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification

Background:Speech patterns have emerged as potential diagnostic markers for conditions with varying etiologies. Machine learning (ML) presents an op...

Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation

Depression is a widespread mental health disorder, and clinical interviews are the gold standard for assessment. However, their reliance on scarce p...

Tackling Hallucination from Conditional Models for Medical Image Reconstruction with DynamicDPS

Hallucinations are spurious structures not present in the ground truth, posing a critical challenge in medical image reconstruction, especially for ...

HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning

In the dynamic landscape of artificial intelligence, the exploration of hallucinations within vision-language (VL) models emerges as a critical fron...

Hybrid Retrieval for Hallucination Mitigation in Large Language Models: A Comparative Analysis

Large Language Models (LLMs) excel in language comprehension and generation but are prone to hallucinations, producing factually incorrect or unsupp...

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the mod...

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffe...

Towards Statistical Factuality Guarantee for Large Vision-Language Models

Advancements in Large Vision-Language Models (LVLMs) have demonstrated promising performance in a variety of vision-language tasks involving image-c...

One-for-More: Continual Diffusion Model for Anomaly Detection

With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also ...

ProAPO: Progressively Automatic Prompt Optimization for Visual Classification

Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their perf...

Human-Centered AI in Multidisciplinary Medical Discussions: Evaluating the Feasibility of a Chat-Based Approach to Case Assessment

In this study, we investigate the feasibility of using a human-centered artificial intelligence (AI) chat platform where medical specialists collabo...

Browse Categories