Investigating Whether Summarizing Raw Clinical Text Affects the Prediction of In-Hospital Violence: A Retrospective Case Study: Évaluation de l'effet du résumé des textes cliniques bruts sur la prédiction de la violence en milieu hospitalier : étude de cas rétrospective.
Journal:
Canadian journal of psychiatry. Revue canadienne de psychiatrie
Published Date:
Aug 31, 2026
Abstract
BackgroundWe evaluated whether summaries and large language models (LLMs) preserve predictive performance for inpatient violence risk and assessed performance across demographics.MethodsWe conducted a retrospective study of inpatient encounters at Harborview Medical Center between March 2021 and September 2023. Cases were inpatient behavioural violence events, and each was matched to 10 control encounters. For each encounter, we used clinical notes from the 72 hours before the event for cases and a matched window for controls. We compared Clinical-Longformer trained on original notes with pipelines that first generated summaries and then classified risk. Summaries were either general or guided by clinician-defined risk entities. We also tested Llama-3.1-8B and MedGemma-27B in zero-shot (ZS) and fine-tuned settings. Primary endpoints were positive-class precision, recall, and F1. We also reported overall and subgroup area under the receiver operating characteristic curve (AUROC) with bootstrap confidence intervals for sex, race, ethnicity, age, and mental health flag.ResultsClinical-Longformer on original notes achieved the strongest performance (F1 = 0.776, AUROC = 0.959). Replacing original notes with general or entity-guided summaries reduced performance (F1 = 0.465 and 0.510). Llama-3.1-8B improved with sequence-classification fine-tuning on original notes but remained below Clinical-Longformer (F1 = 0.701; AUROC = 0.933), while ZS prompting performed poorly. MedGemma-27B also remained inferior, with its best performance from sequence-classification fine-tuning on original notes (F1 = 0.612, AUROC = 0.947). Subgroup AUROCs were high across sex, race, ethnicity, and age, with lower performance among patients with a documented mental health flag.ConclusionDirect classification of original notes remained the most reliable strategy. Summaries of clinical notes may compress or alter cues needed for discrimination, and off-the-shelf prompting was inadequate as a stand-alone predictor. A practical path is to anchor risk scoring in long-context discriminative models and use LLMs for auditable, clinician-facing summaries or rationales. Prospective and external validation are needed before clinical deployment.
Authors
Keywords
No keywords available for this article.