Scale AI
Audio MultiChallenge Evaluation Framework for Dialogue Systems
Pages
33
Time to read
77 mins
Publication
Language
English
Pages
33
Time to read
77 mins
Publication
Language
English
This document is a technical report that introduces Audio MultiChallenge, an open-source benchmark designed to evaluate end-to-end (E2E) spoken dialogue systems in natural multi-turn interactions. It highlights the limitations of existing benchmarks that primarily assess models on synthetic speech and single-turn tasks, emphasizing the need for a more comprehensive evaluation of multi-turn conversational capabilities. The report outlines the framework's structure, which builds upon the text-based MultiChallenge by incorporating new axes such as Voice Editing, which tests robustness to mid-utterance speech repairs, and Audio-Cue challenges that require recalling ambient sounds. The dataset comprises 452 conversations from 47 speakers, with a focus on capturing realistic dialogue patterns and model failures. Analysis reveals that even leading models struggle with the new axes, particularly in maintaining self-coherence and recalling audio cues, indicating significant gaps in current E2E capabilities. The findings suggest a pressing need for improved training and evaluation methods tailored to realistic spoken interactions.