Large Language Models for Individualized Psychoeducational Tools for Psychosis: A Cross-Sectional Study.
Journal:
Early intervention in psychiatry
Published Date:
Aug 1, 2026
Abstract
OBJECTIVE: This study aimed to evaluate the quality of GPT-4-generated responses to commonly asked psychosis-related psychoeducational questions from patients, caregivers and relatives in a first-episode psychosis programme. Evaluation focused on accuracy, clarity, inclusivity, completeness, clinical utility and overall quality. DESIGN: This cross-sectional study employed a qualitative evaluation design. GPT-4, accessed via the ChatGPT interface, generated responses to 20 psychosis-related psychoeducational questions. These questions were developed through iterative discussion and consensus among clinicians working in a first-episode psychosis treatment programme, informed by commonly encountered questions from patients, caregivers and relatives in clinical practice and are provided in Appendix A. The generated responses were subsequently evaluated for their potential clinical applicability. PRIMARY OUTCOME: ChatGPT was presented with 20 psychoeducational questions derived from real-world clinical interactions with patients, caregivers and relatives. Two experts in psychosis independently assessed the responses using a structured six-domain rubric: accuracy (1-3), clarity (1-3), inclusivity (1-3), completeness (0-1), clinical utility (1-5) and overall quality (1-4), where lower scores indicate poorer performance and higher scores indicate stronger performance across domains. Discrepancies in ratings were resolved through discussion and consensus. RESULTS: Using a structured evaluation rubric, LLM-generated responses were assessed across accuracy, clarity, inclusivity, completeness, clinical utility and overall quality. Responses were generally coherent, well-organized and readable across all 20 psychoeducational questions. Performance was strongest in accuracy (M ± SD = 2.88 ± 0.22), clarity (2.93 ± 0.18), completeness (0.93 ± 0.18) and clinical utility (4.35 ± 0.52), indicating that responses were largely correct, understandable and clinically relevant. Inclusivity scores were comparatively lower (2.30 ± 0.41). A descriptive linguistic analysis showed that responses were written at a relatively high Flesch-Kincaid Grade Level (FKGL) (mean = 15.59 ± 1.59), indicating increased reading complexity. Although responses addressed core aspects of the questions, some lacked sufficient nuance for complex or individualized clinical scenarios. CONCLUSIONS: GPT-4, as an example of a large language model (LLM), may have a limited adjunctive role in supporting psychoeducation for psychosis when used within structured and clinician-guided contexts. Although responses were generally readable and clinically relevant, their complexity and variability in inclusivity highlight potential limitations in accessibility for diverse patient populations, and cautious use is warranted given the ongoing concerns regarding accuracy, safety and real-world implementation. Further research is needed before broader clinical integration can be recommended.
Authors
Keywords
No keywords available for this article.