AI Monitoring AI with LLMs: The American College of Radiology's Imaging AI Registry.
Journal:
Journal of the American College of Radiology : JACR
Published Date:
Sep 3, 2026
Abstract
OBJECTIVE: To describe the technical workflow enabling scalable automated artificial intelligence (AI) monitoring in the first national imaging AI registry, Assess-AI. MATERIALS AND METHODS: Large language model (LLM) prompts are developed to extract clinically relevant findings from radiology reports through collaboration between data scientists and subspecialty radiologists. Prompts are optimized using tuning cohorts of use case-specific radiology reports and LLMs available through AWS Bedrock. Such cohorts are used to evaluate prompt accuracy and consistency across 10 repeated runs. Report-AI result pairs submitted to the Assess-AI registry for actively monitored use cases are additionally used to further optimize corresponding prompts. RESULTS: Prompts were developed for nine use cases. In Stage 2 development cohorts, final-prompt agreement with hybrid report-derived reference standard labels was 0.985 for ICH and 0.997 for PE. Because these cohorts informed prompt refinement and label construction, they were not independent validation sets. DISCUSSION: The workflow demonstrates feasible report-finding extraction at scale; independent accuracy and clinical utility remain unestablished. CONCLUSION: LLM-based extraction within Assess-AI enables scalable, report-anchored AI performance monitoring in radiology.
Authors
Keywords
No keywords available for this article.