Generating Patient Documents from Electronic Health Records Using Generative Artificial Intelligence: A Feasibility Study in a Japanese Cancer Center.
Journal:
Applied clinical informatics
Published Date:
Jul 20, 2026
Abstract
OBJECTIVE: Clinical documentation consumes substantial clinician time, potentially detracting from patient care. Generative artificial intelligence (AI) may support drafting discharge summaries and patient referral documents, but feasibility in non-Western-language oncology settings using real-world electronic health record (EHR) data remains insufficiently evaluated. This study assessed feasibility in a Japanese cancer hospital using an enterprise AI system. METHODS: Medical records from 61 consenting adult patients at XXX Cancer Center were analyzed. Although the plan aimed at comprehensive EHR data, actual input was limited to extractable text (physician notes, nursing records); structured laboratory data and imaging, endoscopy, and pathology reports were not directly used, and existing summaries and external referrals were excluded to avoid information leakage. Data were converted to JavaScript Object Notation (JSON); XXX generated 31 discharge summaries and 30 referral documents. Four evaluators scored them; ≥80/100 was an exploratory threshold for draft-level practical utility. Feedback drove one refinement cycle. RESULTS: Generated documents scored approximately 60-70. A score ≥80 was reached by 9 of 31 discharge summaries in each evaluation; for referrals, none reached the threshold initially, whereas 5 of 30 did after refinement. Discharge summary scores did not substantially improve; referral scores did. Raw percent agreement among three non-physician evaluators was high, though chance-corrected agreement varied. Wilcoxon signed-rank tests showed no significant change for discharge summaries (p = 0.866) but significant improvement for referrals (p = 0.006). DISCUSSION AND CONCLUSION: This feasibility study suggests AI may support drafting these documents in a secure environment using real-world Japanese EHR data, though the generated documents did not consistently reach the predefined threshold for draft-level utility. Findings should not be interpreted as demonstrating workload reduction or maximum performance under ideal data conditions. Future studies should evaluate larger datasets, multiple institutions and models, blinded evaluations, actual editing time, clinician acceptance, and workflow impact.
Authors
Keywords
No keywords available for this article.