End-to-End Clinical Validation of a Human-Supervised Large Language Model Agent for Enterprise Surgical Pathology Reporting.
Journal:
Modern pathology : an official journal of the United States and Canadian Academy of Pathology, Inc
Published Date:
Jul 29, 2026
Abstract
Large language models (LLMs) show promise for text-based pathology tasks, yet most reported applications remain experimental, lack formal clinical validation, or operate outside secure, health system-approved environments. We developed and clinically validated a rule-guided, agent-based LLM that assists gastrointestinal (GI) biopsy reporting by automating report structuring while preserving full diagnostic authority with the pathologist. The AI agent (Microsoft 365 Copilot) ran within an enterprise-approved, HIPAA-compliant Microsoft 365 environment, configured with a fixed rule-based system configuration prompt and a quick-text knowledge base. In a prospective validation, 94 GI biopsy cases were evaluated by subspecialty GI pathologists using specimen container labels extracted from the laboratory information system and pathologist-entered shorthand diagnoses. Agent outputs were reviewed for formatting accuracy, organ and procedure identification, shorthand expansion fidelity, blank diagnosis enforcement, and diagnostic safety. The agent preserved specimen part structure and correctly identified organ, sub-organ, and procedure context in 100% of cases; shorthand expansion was accurate in all applicable cases. Minor formatting deviations occurred in 8 cases (8.5%) without affecting diagnostic meaning. Two cases (2%) showed minor diagnostic misinterpretation, in which descriptive container-label terms (e.g., "ulcer," "erosion") were incorporated into diagnostic text; no hallucinated diagnoses were identified. Repeatability testing on cases enriched for descriptive labels showed 81% identical outputs across 105 runs (19% variability), with non-reproducible semantic leakage in 3% of runs. A comparative time study showed faster AI-assisted reporting (mean 39 vs 72 seconds for speech-to-text and 76 seconds for manual typing; ∼33-37 second reductions, p < 0.05), measured across the full workflow through sign-out, with lower variability. By restricting this end-to-end, production-embedded agent to rule-guided structuring, formatting, and controlled shorthand expansion while prohibiting diagnostic inference, the system achieved high efficiency, consistency, and seamless workflow integration on real GI biopsy cases. Low-frequency, stochastic errors and minor variability remain inherent to LLMs despite strict constraints; although infrequent, they indicate such systems are best suited for non-diagnostic, clerical augmentation rather than autonomous use. All output therefore requires pathologist careful review before sign-out. These findings support constrained, agent-based LLMs to safely enhance reporting efficiency while preserving diagnostic responsibility and human oversight.
Authors
Keywords
No keywords available for this article.