Latest AI and machine learning research in schizophrenia for healthcare professionals.
Although multimodal large language models (MLLMs) exhibit remarkable reasoning capabilities on complex multimodal understanding tasks, they still suffer from the notorious hallucination issue: generating outputs misaligned with obvious visual or factual evidence. Currently, training-based solutions, like direct preference optimization (DPO), leverage paired preference data to suppress hallucinat...
We present a perception in reflection paradigm designed to transcend the limitations of current large vision-language models (LVLMs), which are expected yet often fail to achieve perfect perception initially. Specifically, we propose Reflective Perception (RePer), a dual-model reflection mechanism that systematically alternates between policy and critic models, enables iterative refinement of vi...
This paper introduces a comprehensive system for detecting hallucinations in large language model (LLM) outputs in enterprise settings. We present a...
Large Vision-Language Models have demonstrated remarkable performance across various tasks; however, the challenge of hallucinations constrains thei...
Motivated by recent improvements in generative AI and wearable camera devices (e.g. smart glasses and AI-enabled pins), I investigate the ability of...
With the rapid development of 3D printing, the demand for personalized and customized production on the manufacturing line is steadily increasing. E...
Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI)...
This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive ...
The rapid development of multimodal large language models has resulted in remarkable advancements in visual perception and understanding, consolidat...
Multimodal large language models (MLLMs) have demonstrated significant potential in medical Visual Question Answering (VQA). Yet, they remain prone ...
Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. W...
Disorganized thinking is a key diagnostic indicator of schizophrenia-spectrum disorders. Recently, clinical estimates of the severity of disorganize...
The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, rea...
The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability an...
This study addresses the technical bottlenecks in handling long text and the "hallucination" issue caused by insufficient short text information in ...
Composed image retrieval (CIR) enables users to search images using a reference image combined with textual modifications. Recent advances in vision...
Evaluating generative foundation models on open-ended multimodal understanding (MMU) and generation (MMG) tasks across diverse modalities (e.g., ima...
Recent approaches using large-scale pretrained diffusion models for image dehazing improve perceptual quality but often suffer from hallucination is...
Large Language Models (LLMs) frequently generate hallucinated content, posing significant challenges for applications where factuality is crucial. W...
The rapid evolution of Large Vision-Language Models (LVLMs) has highlighted the necessity for comprehensive evaluation frameworks that assess these ...