Ensemble of Domain-Specific Natural Language Processing and Large Language Models for Detecting Suicidal Ideation and Mental Health Conditions in Social Media Text: Development and Evaluation Study.

Journal: JMIR mental health
Published Date:
(1)

Abstract

BACKGROUND: Mental illness contributes substantially to global disability, and public adoption of AI for mental health support is accelerating without commensurate safety evaluation. General-purpose large language models can fail to recognize naturalistic text expressing suicidal ideation, with immediate clinical consequences. Domain-specific natural language processing methods offer a contrasting approach, but ensembles integrating the two have not been formally characterized. OBJECTIVE: We aimed to (1) benchmark 5 modeling approaches for classifying mental health-related social media text, (2) develop and optimize ensembles integrating a domain-specific natural language processing classifier with a fine-tuned large language model, and (3) derive a closed-form framework constraining ensemble weighting in safety-critical classification. METHODS: We analyzed 60,889 publicly available Reddit and Twitter texts spanning 9 categories (anxiety, bipolar disorder, depression, a normal baseline, personality disorder, stress, suicidal ideation, attention-deficit/hyperactivity disorder, and autism spectrum disorder). Labels were derived from the originating subreddit or a self-stated condition and are not clinical diagnoses. Five architectures were compared: prompt-engineered base GPT-4o-mini; fine-tuned GPT-4o-mini; a domain-specific classifier (support vector machine over symptom-informed lexical features); and 2 ensembles of these models, hybrid probability-indicator fusion and soft probability fusion. Data were split into 70%, 10%, and 20% at the post level (n=42,622, n=6089, and n=12,178); weights were selected on the validation split and all reported performance comes from the held-out test split. Closed-form bounds on permissible language model weights were derived for both strategies. CIs are paired bootstrap intervals (n=10,000 replicates); accuracy differences were tested with exact McNemar tests. RESULTS: The domain-specific classifier achieved 91.9% accuracy (95% CI 91.4%-92.4%), exceeding the base (58.7%) and fine-tuned (88.6%) large language models. Both ensembles outperformed either constituent model, reaching 93.9% (soft fusion; 95% CI 93.5%-94.3%) and 93.8% (hybrid fusion; 95% CI 93.4%-94.2%) at a weighting of 60% classifier to 40% fine-tuned model, selected on validation (P<.001 for both against the classifier alone; the 2 ensembles did not differ from each other [P=.16]). The hybrid ensemble reduced the miss rate for the suicidal ideation category from 32.5% (650/2000) to 1.9% (37/2000), a 17.6-fold reduction relative to the base model, although its advantage over the classifier alone within that category was not significant (P=.68). Accuracy fell discontinuously above a 50% language model weight, where the hybrid ensemble reduced exactly to the fine-tuned model; the selected weight satisfies the derived instance-level bound of 44.0%. CONCLUSIONS: Domain-specific clinical grounding remained necessary for stable ensemble performance in this corpus, and probabilistic fusion outperformed both constituent models when language model weight was bounded by a derivable, model-agnostic constraint. Because labels were community-inferred and only one corpus was analyzed, these results are hypotheses requiring external validation on clinically characterized data before use in decision support.

Authors

Keywords

No keywords available for this article.