Fetching the paper…
Reading the bibliography…
Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques.
“Janus-iii: speech-to-speech translation in multiple languages,”
A. Lavie, A. Waibel, L. Levin, M. Finke, D. Gates, M. Gavalda, T. Zeppenfeld, and Z. Puming, · 1997
Earlier work this paper cites.
“Yet another algorithm for pitch tracking,”
K. Kasi and S. Zahorian, · 2002
Earlier work this paper cites.
“The atr multilingual speech-to-speech translation system,”
S. Nakamura, K. Markov, H. Nakaiwa, G. Kikui, H. Kawai, T. Jitsuhiro, J.-S. Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto, · 2006
Earlier work this paper cites.
“Opensmile: the munich versatile and fast open-source audio feature extractor,”
F Eyben, M Wöllmer, and B. Schuller, · 2010
Earlier work this paper cites.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
A. Bérard, O. Pietquin, C. Servan, and L. Besacier, · 2016
Earlier work this paper cites.
“The lj speech dataset,” 2017
I. Keith and J. Linda, · 2017
Earlier work this paper cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Y. Jia, R. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, · 2019
Earlier work this paper cites.
“Unsupervised cross-lingual representation learning for speech recognition,” 2020
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, · 2020
Earlier work this paper cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu, · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“Europarl-st: A multilingual corpus for speech translation of parliamentary debates,”
J. Iranzo-Sánchez, J. Silvestre-Cerda, and J. et al. Jorge, · 2020
Cited alongside, same era.
“Speech resynthesis from discrete disentangled self-supervised representations,”
A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W. N. Hsu, A. Mohamed, and E. Dupoux, · 2021
Cited alongside, same era.
“On generative spoken language modeling from raw audio,”
K. Lakhotia, E. Kharitonov, W. Hsu, Y. Adi, A. Polyak, B. Bolte, T. Nguyen, J. Copet, A. Baevski, and A. Mohamed, · 2021
Cited alongside, same era.
“Textless speech-to-speech translation on real data,”
A. Lee, H. Gong, P. A. Duquenne, H. Schwenk, P. J. Chen, C Wang, S. Popuri, Y. Adi, J. Pino, J. Gu, and W. N. Hsu, · 2022
Later among the works it cites.
“Textless speech emotion conversion using discrete & decomposed representations,”
F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T. Nguyen, M. Rivière, W. Hsu, A. Mohamed, E. Dupoux, and Y. Adi, · 2022
Later among the works it cites.
S. Popuri, P. Chen, and C. et al. Wang, · 2022
Later among the works it cites.
“Speechmatrix: A large-scale mined corpus of multilingual speech-to-speech translations,”
P. Duquenne, H. Gong, and N. et al. Dong, · 2022
Later among the works it cites.
“The flores-101 evaluation benchmark for low-resource and multilingual machine translation,”
N. Goyal, C. Gao, and V. et al. Chaudhary, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A fine-tuned wav2vec 2.0/hubert benchmark for speech emotion recognition, speaker verification and spoken language understanding,”
Y. Wang, A. Boumadane, and A. Heba, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
J. Kim, J. Kong, and J. Son, · 2021
Cited alongside, same era.
“Direct speech-to-speech translation with discrete units,”
A. Lee, P. Chen, C. Wang, J. Gu, S. Popuri, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W. Hsu, · 2022
Cited alongside, same era.
Later among the works it cites.
“Emotional voice conversion: Theory, databases and esd,” 2022
K. Zhou, B. Sisman, R. Liu, and H. Li, · 2022
Later among the works it cites.
“Learning multilingual expressive speech representation for prosody prediction without parallel data,”
J. Duret, Y. Estève, and T. Parcollet, · 2023
Closest in time.
“Fleurs: Few-shot learning evaluation of universal representations of speech,”
A. Conneau, M. Ma, and S. et al. Khanuja, · 2023
Closest in time.