2020

The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Nguyen, Tu Anh, de Seyssel, Maureen, Rozé, Patricia et al.

Understand

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource Speech Benchmark 2021: a suite of 4 black-box, zero-shot metrics probing for the quality of the learned models at 4 linguistic levels: phonetics, lexicon, syntax and semantics.

  • We present the results and analyses of a composite baseline made of the concatenation of three unsupervised systems: self-supervised contrastive representation learning (CPC), clustering (k-means) and language modeling (LSTM or BERT).
  • The language models learn on the basis of the pseudo-text derived from clustering the learned representations.
  • This simple pipeline shows better than chance performance on all four metrics, demonstrating the feasibility of spoken language modeling from raw speech.

Reading the bibliography…