A Controlled Single-Speaker Evaluation of Ambient Voice Technologies' Speech-to-Text Function Using Narrated Surgical Case-Reports.

Journal: Journal of medical systems
Published Date:
(1)

Abstract

Ambient voice technologies (AVT) utilise automatic speech recognition (ASR) and generative artificial intelligence for speech-to-text conversion, and are advocated as efficient, accurate tools for clinical documentation. We investigated (n = 8) AVT systems using surgical case-reports as narratives in a controlled single-speaker evaluation prioritising safety outcomes: specifically 3 commercial transcription-systems (Dragon-Medical-One, Heidi-Health, Tortus); 4 ASR speech-to-text Application Programming Interfaces (Speechmatics-Enhanced, Amazon-Medical-Transcribe, Whisper, GPT4oTranscribe); and an experimental two-stage ASR-Large Language Model (LLM) pipeline incorporating GPT4oTranscribe with GPT-5-LLM generative error correction (GPT4oTranscribe-Corrected-5). Reference-transcripts (n = 100; 32,897 words, range = 44-449, mean = 329/per-transcript) derived from surgical case-reports were recorded and input into AVT systems as audio-recordings for transcription-output generation. Primary outcome was proportion of transcription-outputs containing at least one clinically significant Class 3 error graded for potential harm. Secondary outcomes included transcription accuracy: Domain-Word-Error-Rate (DWER) against SNOMED-CT, lexical-accuracy (ROUGE score) and semantic similarity (BERT, BART scores). To investigate impact of the experimental pipeline on clinically significant errors, raw ASR and LLM-processed transcription-outputs were compared and errors classified as resolved, remaining or newly introduced. Across (n = 800) transcript-outputs, 30-68% (GPT4oTranscribe-Corrected-5; Amazon-Medical-Transcribe and Dragon-Medical-One, respectively) contained at least one clinically significant Class 3 error. LLM-processing reduced the proportion of transcript-outputs affected (53% to 30%), amongst 89 clinically significant errors in ASR-output, 51 were resolved, 38 remained significant and 4 were newly introduced. Reference-transcript length significantly increased odds of Class 3 errors for 4 systems (GPT4oTranscribe, Tortus, Amazon-Medical-Transcribe, Dragon-Medical-One; odds ratio 2.65-3.34/additional 100 reference-words). Errors with potential to cause severe harm or death (NHS England Level 3-4) totalled (n = 53) across systems, concentrated in the domains of medication-type/dose (39.6%), and investigations/laboratory results (30.2%). Performance varied across systems for transcription metrics (P < 0.001). GPT4oTranscribe-Corrected-5 (DWER = 3.60%) and Heidi-Health (DWER = 5.67%) performed best for medical terminology; Amazon-Medical-Transcribe worst (DWER = 24.03%). For lexical and semantic similarity GPT4oTranscribe-Corrected-5 performed best, followed by Heidi-Health. The ASR-LLM pipeline reduced clinically significant errors but did introduce new ones in 3% of transcript-outputs. The safety and effectiveness of these systems in clinical practice requires further evaluation.

Authors

Keywords

No keywords available for this article.