Fetching the paper…
Reading the bibliography…
In cross-lingual speech synthesis, the speech in various languages can be synthesized for a monoglot speaker.
1929
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2Wav: End-to-end speech synthesis,” in Proceedings of International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman et al. , “Deep voice: Real-time neural text-to-speech,” in Proceedings of International Conference on Machine Learning (ICML) , 2017, pp. 195–204
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Proceedings of Interspeech , 2017, pp. 2616–2620. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2017-950
2017
Earlier work this paper cites.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “Voiceloop: Voice fitting and synthesis via a phonological loop,” in Proceedings of International Conference on Learning Representations (ICLR) , 2018
2018
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” in Proceedings of Interspeech , 2018, pp. 4006–4010
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proceedings of ICASSP . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
P. Baljekar, S. K. Rallabandi, and A. W. Black, “An investigation of convolution attention based models for multilingual speech synthesis of Indian languages.” in Proceedings of Interspeech , 2018, pp. 2474–2478
2018
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in Proceedings of ICASSP . IEEE, 2018, pp. 5329–5333
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 6706–6713
2019
Cited alongside, same era.
Z. Liu and B. Mak, “Cross-lingual multi-speaker text-to-speech synthesis for voice cloning without using parallel corpus for unseen speakers,” 2019
2019
Later among the works it cites.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding,” in Proceedings of Interspeech , 2019, pp. 2105–2109
2019
Later among the works it cites.
D. Duckworth, A. Neelakantan, B. Goodrich, L. Kaiser, and S. Bengio, “Parallel scheduled sampling,” 2019
2019
Later among the works it cites.
P. Zhou, R. Fan, W. Chen, and J. Jia, “Improving generalization of transformer for speech recognition with parallel schedule sampling and relative positional embedding,” 2019
2019
Later among the works it cites.
J. Yang and L. He, “Towards universal text-to-speech,” in Proceedings of Interspeech , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Li, Y. Zhang, T. Sainath, Y. Wu, and W. Chan, “Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,” in Proceedings of ICASSP . IEEE, 2019, pp. 5621–5625
2019
Cited alongside, same era.
2019
Cited alongside, same era.
E. Nachmani and L. Wolf, “Unsupervised polyglot text-to-speech,” in Proceedings of ICASSP . IEEE, 2019, pp. 7055–7059
2019
Cited alongside, same era.
2020
Later among the works it cites.
Z. Cai, Y. Yang, and M. Li, “Cross-lingual multispeaker text-to-speech under limited-data scenario,” 2020
2020
Later among the works it cites.
L. Chen, K. Lee, L. He, and F. Soong, “On early-stop clustering for speaker diarization,” in Proceedings Odyssey 2020 the Speaker and Language Recognition Workshop , 2020, pp. 110–116
2020
Later among the works it cites.