mpathic
Benchmarking Speech AI for Speaker Attribution in Clinical Settings
Pages
1
Time to read
8 mins
Publication
Language
English
Pages
1
Time to read
8 mins
Publication
Language
English
This technical report benchmarks commercial and open-source Automatic Speech Recognition (ASR) systems for speaker attribution in real-world clinical conversations. It addresses the critical need for accurate speaker identification beyond traditional role labels, which are insufficient for clinical fidelity and safety. The report introduces cpHEWER, a novel evaluation metric that emphasizes clinically significant transcription errors, particularly in multi-speaker dialogues. The study evaluates three ASR systems—Rev AI, Assembly AI, and Whisper + Pyannote—using the AnnoMI dataset, which simulates diverse clinical audio conditions. Findings indicate that existing metrics may underestimate the impact of attribution errors, especially in the presence of accent variation. Recommendations for improving ASR performance include developing compound error metrics, enhancing speaker count detection accuracy, and expanding diverse clinical datasets. The report emphasizes the importance of context-aware validation to ensure safe and effective ASR deployment in clinical trials.