2021

Direct Simultaneous Speech-to-Speech Translation with Variational Monotonic Multihead Attention

Ma, Xutai, Gong, Hongyu, Liu, Danni et al.

Understand

We present a direct simultaneous speech-to-speech translation (Simul-S2ST) model, Furthermore, the generation of translation is independent from intermediate text representations.

  • Our approach leverages recent progress on direct speech-to-speech translation with discrete units, in which a sequence of discrete representations, instead of continuous spectrogram features, learned in an unsupervised manner, are predicted from the model and passed directly to a vocoder for speech synthesis on-the-fly.
  • We also introduce the variational monotonic multihead attention (V-MMA), to handle the challenge of inefficient policy learning in speech simultaneous translation.
  • The simultaneous policy then operates on source speech features and target discrete units.

Reading the bibliography…