Latest AI and machine learning research in schizophrenia for healthcare professionals.
Foundation Models that are capable of processing and generating multi-modal data have transformed AI's role in medicine. However, a key limitation of their reliability is hallucination, where inaccurate or fabricated information can impact clinical decisions and patient safety. We define medical hallucination as any instance in which a model generates misleading medical content. This paper exami...
Vision-language models in pathology enable multimodal case retrieval and automated report generation. Many of the models developed so far, however, have been trained on pathology reports that include information which cannot be inferred from paired whole slide images (e.g., patient history), potentially leading to hallucinated sentences in generated reports. To this end, we investigate how the s...
Self-supervised learning (SSL) vision encoders learn high-quality image representations and thus have become a vital part of developing vision modal...
Capturing subtle speech disruptions across the psychosis spectrum is challenging because of the inherent variability in speech patterns. This variab...
Large Vision-Language Models (LVLMs) integrate image encoders with Large Language Models (LLMs) to process multi-modal inputs and perform complex vi...
Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications....
Large Language Models (LLMs) are prone to hallucinations, e.g., factually incorrect information, in their responses. These hallucinations present ch...
Multimodal Large Language Models (MLLMs) have shown impressive performance in vision and text tasks. However, hallucination remains a major challeng...
Advancements in Large Language Models (LLMs) and their increasing use in medical question-answering necessitate rigorous evaluation of their reliabi...
Vision language models (VLM) demonstrate sophisticated multimodal reasoning yet are prone to hallucination when confronted with knowledge conflicts,...
Despite recent advances in Novel View Synthesis (NVS), generating high-fidelity views from single or sparse observations remains a significant chall...
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, p...
Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple s...
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders the...
Large Language Models (LLMs) are transforming healthcare through the development of LLM-based agents that can understand, reason about, and assist w...
Recent advances in large language models (LLMs) have shown promising improvements, often surpassing existing methods across a wide range of downstre...
Multimodal Large Language Models (MLLMs) represent the cutting edge of AI technology, with DeepSeek models emerging as a leading open-source alterna...
Recent advancements in large language models have demonstrated significant potential in the automated construction of knowledge graphs from unstruct...
This study investigates the potential of multimodal data integration, which combines electroencephalogram (EEG) data with sociodemographic character...
Large Vision-Language Models (LVLMs) exhibit impressive multimodal reasoning capabilities but remain highly susceptible to object hallucination, whe...