Modern generative large language models in epilepsy care: a scoping review of current applications, challenges, and future directions.
Journal:
Acta epileptologica
Published Date:
Sep 1, 2026
Abstract
Modern generative large language models (LLMs) are increasingly being evaluated in epilepsy-related clinical tasks, but the evidence remains fragmented and their safe clinical role is uncertain. We conducted a scoping review following the PRISMA-ScR framework, searching PubMed, Embase, and the Web of Science Core Collection from inception to May 2026, to map current applications of modern generative LLMs in epilepsy-related clinical contexts and identify their reported benefits, limitations, and research gaps. Eligible studies evaluated generative LLMs or generative foundation models as the primary analytical or assistive engine in clinical or clinically oriented epilepsy tasks; non-generative natural language processing, conventional machine-learning, and signal-modeling studies served as contextual comparators only. Two reviewers independently screened records and extracted data for descriptive mapping and narrative synthesis. Twenty-four studies were included. Evidence was relatively more developed for text-centered tasks, including information extraction from electronic health records, question answering, and documentation support, while applications in differential diagnosis, presurgical evaluation, prognostic assessment, and patient communication remained early-stage. Across studies, the evidence base was heterogeneous, with frequent reliance on retrospective or simulated designs, single-center data, and limited external validation, alongside persistent concerns about hallucination, bias, model opacity, and workflow integration. Current evidence suggests a limited role for cautious, clinician-supervised use of generative LLMs in selected text-heavy epilepsy tasks, particularly extraction, summarization, documentation support, and patient education, but falls short of justifying independent or routine clinical deployment. Future studies should use narrower task definitions, epilepsy-specific datasets, transparent reporting of model versions and prompt design, external validation, and explicit documentation of clinician verification workflows.
Authors
Keywords
No keywords available for this article.