2023

AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data

Yu, Jianwei, Chen, Hangting, Bian, Yanyao et al.

Understand

Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs).

  • However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained.
  • One reason for this limitation is that a considerable amount of the publicly available speech data is compromised by background noise, speech overlapping, lack of speech segmentation information, missing speaker labels, and incomplete transcriptions, which can largely hinder their usefulness.
  • On the other hand, human annotation of speech data is both time-consuming and costly.

Reading the bibliography…