The use of large language models in automated depression detection.
Journal:
Acta psychologica
Published Date:
Aug 11, 2026
Abstract
BACKGROUND: Large language models have been evaluated on many healthcare tasks, including depression screening. However, it is unclear whether estimates of performance are accurate, especially in a setting with realistic clinical constraints. METHODS: We use the publicly-available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) to give performance estimates of LLMs that respect patient privacy and could be feasibly deployed in a clinical setting. These models are locally run and under 15 billion parameters. RESULTS: Accuracy, sensitivity, and specificity of the models we evaluated ranged from 0.233-0.677, 0.041-0.929, 0.000-0.729, respectively. There are significant differences in the performance we observed versus other studies that evaluate commercial models. We also demonstrate poor agreement amongst different LLMs. CONCLUSION: Current performance estimates of LLMs with respect to depression screening are most likely optimistic. When restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.
Authors
Keywords
No keywords available for this article.