A Secure User Interface for Preclinical Evaluation of AI in Patient Portal Message Management: Tutorial.
Journal:
JMIR medical informatics
Published Date:
Jul 20, 2026
Abstract
The growing use of AI to support patient portal message management requires rigorous preclinical evaluation. Directly testing AI within electronic health record (EHR) systems poses significant safety, workflow, and data-governance risks. Here, we present a technical feasibility report on a secure user interface (UI) sandbox designed to enable clinical and technical teams to experiment with AI for portal messaging before clinical integration. In this context, a "sandbox" refers to a controlled, nonproduction environment that allows safe testing, prompt iteration, and evaluation of AI outputs without impacting live EHR systems or patient care. We developed a web UI in Python 3 with a modular backend for data handling and AI task execution that operates entirely within the institutional firewall. The system runs in a secure research environment equipped with an NVIDIA GRID T4-1Q graphics processing unit (GPU) and institutional access controls. We designed a deidentification pipeline to remove or replace personal health identifiers and assessed its precision. The platform supports single-message and batch workflows and exposes example large language model (LLM)-enabled tasks such as authorship identification, message categorization, criticality flagging, and response drafting using zero-shot, one-shot, and few-shot prompting. The system successfully executed end-to-end workflows to ingest messages, run individual or batch AI analyses, and present outputs for review. Personal health information partial masking was applied across the corpus using a deidentification pipeline validated against 110 manually adjudicated entities (sensitivity 95.1%, precision 82.1%). We ran use cases with an institutional review board-approved corpus of a dementia-relevant subset of 6941 patient portal messages categorized as "medical advice requests" from 497 unique patients. With the support of the UI, we tested which prompting strategies yielded interpretable outputs for authorship identification, categorization, and criticality flagging, and whether response drafting produced editable clinician starting points. A token-based cost readout provided transparent operating estimates for LLM-backed tasks. This framework offers a practical, secure path to test AI behavior on real messages without affecting live EHR workflows and thus supports exploratory testing, prompt iteration, and comparative analyses, including LLM prompts versus baseline models, while preserving governance boundaries. We discuss design choices, safety controls, and the limits of a sandbox approach. A secure, UI-based sandbox enables health system teams to evaluate AI for patient portal messaging before clinical integration. The goal is not to assume benefit but to generate evidence about feasibility, risks, and fit to clinical needs in a controlled setting.
Authors
Keywords
No keywords available for this article.