Fetching the paper…
Reading the bibliography…
YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Proc. Interspeech 2018 , 2018, pp. 1086–1090. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1929
1929
Earlier work this paper cites.
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer, “Crowdmos: An approach for crowdsourcing mean opinion score studies,” in Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on . IEEE, 2011, pp. 2416–2419
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” University of Edinburgh. The Centre for Speech Technology Research (CSTR) , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real NVP,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings . OpenReview.net, 2017. [Online]. Available: https://openreview.net/forum?id=HkpbnH9lx
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Ito et al. , “The lj speech dataset,” 2017
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” 2017
2017
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in neural information processing systems , 2018, pp. 4480–4490
2018
Earlier work this paper cites.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in Advances in Neural Information Processing Systems , 2018, pp. 10 019–10 029
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4879–4883
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5329–5333
2018
Earlier work this paper cites.
Y. Cao, X. Wu, S. Liu, J. Yu, X. Li, Z. Wu, X. Liu, and H. Meng, “End-to-end code-switched tts with mix of monolingual recordings,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6935–6939
2019
Cited alongside, same era.
Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Z. Chen, R. Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, “Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,” Proc. Interspeech 2019 , pp. 2080–2084, 2019
2019
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Cited alongside, same era.
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, “High fidelity speech synthesis with adversarial networks,” in International Conference on Learning Representations , 2019
2020
Later among the works it cites.
J. S. Chung, J. Huh, S. Mun, M. Lee, H. S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, “In defence of metric learning for speaker recognition,” in Interspeech , 2020
2020
Later among the works it cites.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020 , H. Meng, B. Xu, and T. F. Zheng, Eds. ISCA, 2020, pp. 2757–2761. [Online]. Available: https://doi.org/10.21437/Interspeech.2020-2826
2020
Later among the works it cites.
E. Casanova, A. C. Junior, C. Shulby, F. S. de Oliveira, J. P. Teixeira, M. A. Ponti, and S. M. Aluisio, “Tts-portuguese corpus: a corpus for speech synthesis in brazilian portuguese,” 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Munich Artificial Intelligence Laboratories GmbH, “The m-ailabs speech dataset – caito,” 2017. [Online]. Available: https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/
2019
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, pp. 8026–8037, 2019
2019
Cited alongside, same era.
C. Jemine, “Master thesis: Real-time voice cloning,” 2019
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “Autovc: Zero-shot voice style transfer with only autoencoder loss,” in International Conference on Machine Learning . PMLR, 2019, pp. 5210–5219
2019
Cited alongside, same era.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6184–6188
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Cited alongside, same era.
2020
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the 12th Language Resources and Evaluation Conference , 2020, pp. 4218–4222
2020
Later among the works it cites.
2020
Later among the works it cites.
E. Casanova, C. Shulby, E. Gölge, N. M. Müller, F. S. de Oliveira, A. Candido Jr., A. da Silva Soares, S. M. Aluisio, and M. A. Ponti, “SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech Model,” in Proc. Interspeech 2021 , 2021, pp. 3645–3649
2021
Closest in time.
N. Kumar, S. Goel, A. Narang, and B. Lall, “Normalization Driven Zero-Shot Multi-Speaker Speech Synthesis,” in Proc. Interspeech 2021 , 2021, pp. 1354–1358
2021
Closest in time.
2021
Closest in time.
S. Li, B. Ouyang, L. Li, and Q. Hong, “Light-tts: Lightweight multi-speaker multi-lingual text-to-speech,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 8383–8387
2021
Closest in time.
2021
Closest in time.
D. Xin, Y. Saito, S. Takamichi, T. Koriyama, and H. Saruwatari, “Cross-Lingual Speaker Adaptation Using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis,” in Proc. Interspeech 2021 , 2021, pp. 1614–1618
2021
Closest in time.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021
2021
Closest in time.
X. Hao, X. Su, R. Horaud, and X. Li, “Fullsubnet: A full-band and sub-band fusion model for real-time single-channel speech enhancement,” ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Jun 2021. [Online]. Available: http://dx.doi.org/10.1109/ICASSP39728.2021.9414177
2021
Closest in time.