The use of large language models in automated depression detection.

Journal: Acta psychologica
Published Date:

Abstract

BACKGROUND: Large language models have been evaluated on many healthcare tasks, including depression screening. However, it is unclear whether estimates of performance are accurate, especially in a setting with realistic clinical constraints. METHODS: We use the publicly-available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) to give performance estimates of LLMs that respect patient privacy and could be feasibly deployed in a clinical setting. These models are locally run and under 15 billion parameters. RESULTS: Accuracy, sensitivity, and specificity of the models we evaluated ranged from 0.233-0.677, 0.041-0.929, 0.000-0.729, respectively. There are significant differences in the performance we observed versus other studies that evaluate commercial models. We also demonstrate poor agreement amongst different LLMs. CONCLUSION: Current performance estimates of LLMs with respect to depression screening are most likely optimistic. When restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.

Authors

Keywords

No keywords available for this article.