Evaluating Artificial Intelligence Translation Tools for Language Equivalence of Oncology-Informed Consent Forms From English to Spanish.

Journal: JCO oncology practice
Published Date:

Abstract

PURPOSE: Approximately 8% of the US population speaks primary languages other than English. Limited English proficiency (LEP) contributes to under-representation of Hispanic patients in oncology clinical trials. Although certified translation services exist, they are time-consuming and costly. Artificial intelligence (AI)-generated translations of informed consent forms (ICFs) could provide low-cost alternatives, but data on accuracy and safety remain limited. We evaluated language equivalence of English-to-Spanish translations for three oncology clinical trial ICFs using two general-purpose AI translation tools (DeepL Pro and ChatGPT-4o) and a medically trained AI translation tool (Med_English2Spanish) compared with certified translations. METHODS: Translational equivalence was assessed using a five-point Likert scale on five domains: Semantic, Idiomatic, Experiential, Conceptual, and Safety. Two native Spanish-speaking bilingual board-certified physicians independently scored each translation. Weighted Cohen's kappa determined inter-rater reliability, and the two-sample t-test compared AI-generated and certified translations. RESULTS: Weighted Cohen's kappa (0.95, 95% CI 0.85 to 0.97) exhibited high inter-rater agreement. Certified translations exhibited the highest equivalence (mean = 4.99, SD = 0.02). ChatGPT-4o similarly demonstrated high equivalence (mean = 4.89, SD = 0.17). DeepL Pro scored well (mean = 4.43, SD = 0.07) but lower than certified translation (P < 0.001). Med_English2Spanish demonstrated the lowest degree of equivalence (mean = 3.32, SD = 0.40) compared with certified translations (P < 0.001). CONCLUSION: Low-cost AI translations of ICFs exhibited variable language equivalence compared with certified translations across several domains. ChatGPT-4o scored nearly equivalent across domains in translating procedural trial information. While AI-generated translations are currently not suitable for clinical deployment without human review, this exploratory study supports further analysis of AI translation tools for reducing language barriers to LEP population enrollment.

Authors

Keywords

No keywords available for this article.