Fetching the paper…
Reading the bibliography…
Speech-to-speech translation (S2ST) enables spoken communication between people talking in different languages.
S. Nakamura, K. Markov, H. Nakaiwa, G. Kikui, H. Kawai, T. Jitsuhiro, J. Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto, “The ATR multilingual speech-to-speech translation system,” IEEE Trans. Speech Audio Process. , vol. 14, no. 2, pp. 365–376, 2006
2006
Earlier work this paper cites.
R. Aharoni, M. Johnson, and O. Firat, “Massively multilingual neural machine translation,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, 2019, pp. 3874–3884
2019
Earlier work this paper cites.
K. Park and T. Mulc, “CSS10: A collection of single speaker speech datasets for 10 languages,” in Interspeech 2019, 20th Annual Conference of the International Speech Communication Association, Graz, Austria, 15-19 September 2019 . ISCA, 2019, pp. 1566–1570
2019
Earlier work this paper cites.
J. Iranzo-Sánchez, J. A. Silvestre-Cerdà, J. Jorge, N. Roselló, A. Giménez, A. Sanchís, J. Civera, and A. Juan, “Europarl-st: A multilingual corpus for speech translation of parliamentary debates,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020 . IEEE, 2020, pp. 8229–8233
2020
Earlier work this paper cites.
V. Pratap, A. Sriram, P. Tomasello, A. Y. Hannun, V. Liptchinsky, G. Synnaeve, and R. Collobert, “Massively multilingual ASR: 50 languages, 1 model, 1 billion parameters,” in Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020 . ISCA, 2020, pp. 4751–4755
2020
Earlier work this paper cites.
T. Nekvinda and O. Dusek, “One model, many languages: Meta-learning for multilingual text-to-speech,” in Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020 . ISCA, 2020, pp. 2972–2976
2020
Earlier work this paper cites.
Z. Wang, Z. C. Lipton, and Y. Tsvetkov, “On negative interference in multilingual models: Findings and A meta-learning treatment,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 . Association for Computational Linguistics, 2020, pp. 4438–4450
2020
Earlier work this paper cites.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Online: Association for Computational Linguistics, Oct. 2020, pp. 38–45. [Online]. Available: https://www.aclweb.org/anthology/2020.emnlp-demos.6
2020
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020 . European Language Resources Association, 2020, pp. 4218–4222
2020
Cited alongside, same era.
C. Wang, M. Rivière, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. M. Pino, and E. Dupoux, “Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 . Association for Computational Linguistics, 2021, pp. 993–1003
2021
Cited alongside, same era.
Y. Tang, H. Gong, X. Li, C. Wang, J. M. Pino, H. Schwenk, and N. Goyal, “FST: the FAIR speech translation system for the IWSLT21 multilingual shared task,” in Proceedings of the 18th International Conference on Spoken Language Translation, IWSLT 2021, Bangkok, Thailand (online), August 5-6, 2021 . Association for Computational Linguistics, 2021, pp. 131–137
Y. Jia, M. T. Ramanovich, Q. Wang, and H. Zen, “CVSS corpus and massively multilingual speech-to-speech translation,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, LREC 2022, Marseille, France, 20-25 June 2022 . European Language Resources Association, 2022, pp. 6691–6703
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “FLEURS: few-shot learning evaluation of universal representations of speech,” in IEEE Spoken Language Technology Workshop, SLT 2022, Doha, Qatar, January 9-12, 2023 . IEEE, 2022, pp. 798–805
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W. Hsu, A. Mohamed, and E. Dupoux, “Speech resynthesis from discrete disentangled self-supervised representations,” in Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, 30 August - 3 September 2021 . ISCA, 2021, pp. 3615–3619
2021
Cited alongside, same era.
Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, “Translatotron 2: High-quality direct speech-to-speech translation with voice preservation,” in International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA , ser. Proceedings of Machine Learning Research, vol. 162. PMLR, 2022, pp. 10 120–10 134
2022
Cited alongside, same era.
A. Lee, H. Gong, P. Duquenne, H. Schwenk, P. Chen, C. Wang, S. Popuri, Y. Adi, J. M. Pino, J. Gu, and W. Hsu, “Textless speech-to-speech translation on real data,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022 . Association for Computational Linguistics, 2022, pp. 860–872
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Later among the works it cites.
S. Popuri, P. Chen, C. Wang, J. Pino, Y. Adi, J. Gu, W. Hsu, and A. Lee, “Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation,” in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022 . ISCA, 2022, pp. 5195–5199. [Online]. Available: https://doi.org/10.21437/Interspeech.2022-11032
2022
Later among the works it cites.
2022
Later among the works it cites.
X.-P. Nguyen, S. Popuri, C. Wang, Y. Tang, I. Kulikov, and H. Gong, “Improving speech-to-speech translation through unlabeled text,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.