2020

Multi-task self-supervised learning for Robust Speech Recognition

Ravanelli, Mirco, Zhong, Jianyuan, Pascual, Santiago et al.

Understand

Despite the growing interest in unsupervised learning, extracting meaningful knowledge from unlabelled audio remains an open challenge.

  • To take a step in this direction, we recently proposed a problem-agnostic speech encoder (PASE), that combines a convolutional encoder followed by multiple neural networks, called workers, tasked to solve self-supervised problems (i.e., ones that do not require manual annotations as ground truth).
  • PASE was shown to capture relevant speech information, including speaker voice-print and phonemes.
  • This paper proposes PASE+, an improved version of PASE for robust speech recognition in noisy and reverberant environments.

Reading the bibliography…