Latest AI and machine learning research in schizophrenia for healthcare professionals.
Capturing subtle speech disruptions across the psychosis spectrum is challenging because of the inherent variability in speech patterns. This variability reflects individual differences and the fluctuating nature of symptoms in both clinical and non-clinical populations. Accounting for uncertainty in speech data is essential for predicting symptom severity and improving diagnostic precision. Spe...
Large Vision-Language Models (LVLMs) integrate image encoders with Large Language Models (LLMs) to process multi-modal inputs and perform complex visual tasks. However, they often generate hallucinations by describing non-existent objects or attributes, compromising their reliability. This study analyzes hallucination patterns in image captioning, showing that not all tokens in the generation pr...
Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications....
Large Language Models (LLMs) are prone to hallucinations, e.g., factually incorrect information, in their responses. These hallucinations present ch...
Multimodal Large Language Models (MLLMs) have shown impressive performance in vision and text tasks. However, hallucination remains a major challeng...
Advancements in Large Language Models (LLMs) and their increasing use in medical question-answering necessitate rigorous evaluation of their reliabi...
Vision language models (VLM) demonstrate sophisticated multimodal reasoning yet are prone to hallucination when confronted with knowledge conflicts,...
Despite recent advances in Novel View Synthesis (NVS), generating high-fidelity views from single or sparse observations remains a significant chall...
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, p...
Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple s...
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders the...
Large Language Models (LLMs) are transforming healthcare through the development of LLM-based agents that can understand, reason about, and assist w...
Recent advances in large language models (LLMs) have shown promising improvements, often surpassing existing methods across a wide range of downstre...
Multimodal Large Language Models (MLLMs) represent the cutting edge of AI technology, with DeepSeek models emerging as a leading open-source alterna...
Recent advancements in large language models have demonstrated significant potential in the automated construction of knowledge graphs from unstruct...
This study investigates the potential of multimodal data integration, which combines electroencephalogram (EEG) data with sociodemographic character...
Large Vision-Language Models (LVLMs) exhibit impressive multimodal reasoning capabilities but remain highly susceptible to object hallucination, whe...
Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing model...
Magnetic Resonance Imaging generally requires long exposure times, while being sensitive to patient motion, resulting in artifacts in the acquired i...
Hallucination has been a long-standing and inevitable problem that hinders the application of Large Vision-Language Models (LVLMs) in domains that r...