Ricoh
Self-Supervised Learning for ASR with Data Augmentation
Pages
9
Time to read
18 mins
Publication
Language
English
Pages
9
Time to read
18 mins
Publication
Language
English
This technical report presents a novel approach to self-supervised learning for automatic speech recognition (ASR) that includes methods for generating self-supervised labels and data augmentation techniques. The report outlines the challenges associated with utilizing large-scale pre-trained models in environments with limited data, particularly the high costs of supervised learning for transcription. It discusses the effectiveness of self-supervised learning in ASR tasks and proposes a new method that uniquely determines target labels using a deterministic algorithm. Experimental results demonstrate that the proposed method improves recognition rates compared to non pre-trained models under limited data conditions. Additionally, the report details a data augmentation strategy that enhances model robustness while retaining critical speech characteristics. This approach aims to facilitate the development of ASR systems that require less transcribed speech data, thereby addressing the limitations of traditional methods.