2020

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Wang, Changhan, Wu, Anne, Pino, Juan

Understand

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets.

  • Nevertheless, current datasets cover a limited number of languages.
  • With the aim to foster research in massive multilingual speech translation and speech translation for low resource language pairs, we release CoVoST 2, a large-scale multilingual speech translation corpus covering translations from 21 languages into English and from English into 15 languages.
  • This represents the largest open dataset available to date from total volume and language coverage perspective.

Reading the bibliography…