Fetching the paper…
Reading the bibliography…
Speech-to-speech translation (S2ST) converts input speech to speech in another language.
W. E. Cooper and M. Danly, “Segmental and temporal aspects of utterance-final lengthening,” Phonetica , vol. 38, no. 1-3, pp. 106–115, 1981
1981
Earlier work this paper cites.
R. Berkovits, “Utterance-final lengthening and the duration of final-stop closures,” Journal of Phonetics , vol. 21, no. 4, pp. 479–489, 1993
1993
Earlier work this paper cites.
A. Lavie, A. Waibel, L. S. Levin, M. Finke, D. Gates, M. Gavaldà, T. Zeppenfeld, and P. Zhan, “Janus-iii: speech-to-speech translation in multiple languages,” in Proc. ICASSP , 1997
1997
Earlier work this paper cites.
S. Nakamura, K. Markov, H. Nakaiwa, G. Kikui, H. Kawai, T. Jitsuhiro, J. Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto, “The ATR multilingual speech-to-speech translation system,” IEEE Trans. Speech Audio Process. , 2006
2006
Earlier work this paper cites.
J. Kominek, T. Schultz, and A. W. Black, “Synthesizer voice quality of new languages calibrated with mean mel cepstral distortion,” in Proc. SLTU , 2008
2008
Earlier work this paper cites.
M. Post, G. Kumar, A. Lopez, D. G. Karakos, C. Callison-Burch, and S. Khudanpur, “Improved speech-to-text translation with the fisher and callhome spanish-english speech translation corpus,” in Proc. IWSLT , 2013
2013
Earlier work this paper cites.
T. Schultz and T. Schlippe, “GlobalPhone: Pronunciation dictionaries in 20 languages,” in Proc LREC , 2014
2014
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proc. ACL , 2016
2016
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. V. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ Speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi,” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on Mel spectrogram predictions,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proc. EMNLP , 2018
2018
Earlier work this paper cites.
M. Post, “A call for clarity in reporting BLEU scores,” in Proc. WMT , 2018
2018
Earlier work this paper cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Speech-to-speech translation between untranscribed unknown languages,” in Proc. ASRU , 2019
2019
Cited alongside, same era.
Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, “Direct speech-to-speech translation with a sequence-to-sequence model,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
M. Ma, L. Huang, H. Xiong, R. Zheng, K. Liu, B. Zheng, C. Zhang, Z. He, H. Liu, X. Li, H. Wu, and H. Wang, “STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework,” in Proc. ACL , 2019
2019
Cited alongside, same era.
T. Yanagita, S. Sakti, and S. Nakamura, “Neural iTTS: Toward Synthesizing Speech in Real-time with End-to-end Neural Text-to-Speech Framework,” in Proc. ISCA SSW , 2019
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS , 2020
2020
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proc. LREC , Marseille, France, 2020
2020
Later among the works it cites.
R. Fukuda, S. Novitasari, Y. Oka, Y. Kano, Y. Yano, Y. Ko, H. Tokuyama, K. Doi, T. Yanagita, S. Sakti, K. Sudoh, and S. Nakamura, “Simultaneous speech-to-speech translation system with transformer-based incremental asr, mt, and tts,” in Proc. O-COCOSDA , 2021
2021
Closest in time.
B. Stephenson, T. Hueber, L. Girin, and L. Besacier, “Alternate Endings: Improving Prosody for Incremental Neural TTS with Predicted Future Text Input,” in Proc. Interspeech , 2021
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in Proc. NAACL , 2019
2019
Cited alongside, same era.
K. Park and T. Mulc, “CSS10: A collection of single speaker speech datasets for 10 languages,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
X. Ma, J. Pino, and P. Koehn, “SimulMT to SimulST: Adapting simultaneous text translation to end-to-end simultaneous speech translation,” in Proc. AACL , 2020
2020
Cited alongside, same era.
M. Ma, B. Zheng, K. Liu, R. Zheng, H. Liu, K. Peng, K. Church, and L. Huang, “Incremental text-to-speech synthesis with prefix-to-prefix framework,” in Proc. EMNLP , 2020
2020
Cited alongside, same era.
B. Stephenson, L. Besacier, L. Girin, and T. Hueber, “What the future brings: Investigating the impact of lookahead for incremental neural TTS,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
D. S. R. Mohan, R. Lenain, L. Foglianti, T. H. Teh, M. Staib, A. Torresquintero, and J. Gao, “Incremental text to speech for neural sequence-to-sequence models using reinforcement learning,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, “Fairseq S2T: Fast speech-to-text modeling with Fairseq,” in Proc. AACL , 2020
2020
Cited alongside, same era.
T. Saeki, S. Takamichi, and H. Saruwatari, “Incremental text-to-speech synthesis using pseudo lookahead with large pretrained language model,” IEEE Signal Process. Lett. , vol. 28, 2021
2021
Closest in time.
T. Saeki, S. Takamichi, and Saruwatari, “Low-latency incremental text-to-speech synthesis with distilled context prediction network,” in Proc. ASRU , 2021
2021
Closest in time.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2021
2021
Closest in time.
R. J. Weiss, R. J. Skerry-Ryan, E. Battenberg, S. Mariooryad, and D. P. Kingma, “Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis,” in Proc. ICASSP , 2021
2021
Closest in time.
C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. Pino, “fairseq sˆ2: A scalable and integrable speech synthesis toolkit,” in Proc. EMNLP , 2021
2021
Closest in time.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised cross-lingual representation learning for speech recognition,” in Proc. Interspeech , 2021
2021
Closest in time.
A. Lee, P.-J. Chen, C. Wang, J. Gu, S. Popuri, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W.-N. Hsu, “Direct speech-to-speech translation with discrete units,” in Proc. ACL , 2022
2022
Closest in time.
Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, “Translatotron 2: Robust direct speech-to-speech translation,” in Proc. ICML , 2022
2022
Closest in time.