Fetching the paper…
Reading the bibliography…
Training a multi-speaker Text-to-Speech (TTS) model from scratch is computationally expensive and adding new speakers to the dataset requires the model to be re-trained.
R. M. French, “Catastrophic forgetting in connectionist networks,” Trends in Cognitive Sciences , vol. 3, no. 4, pp. 128–135, 1999. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1364661399012942
1999
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, http://www.deeplearningbook.org
2016
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell, “Overcoming catastrophic forgetting in neural networks,” Proceedings of the National Academy of Sciences , vol. 114, no. 13, pp. 3521–3526, 2017. [Online]. Available: https://www.pnas.org/content/114/13/3521
2017
Earlier work this paper cites.
S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert, “icarl: Incremental classifier and representation learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , July 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: 2000-speaker neural text-to-speech,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=HJtEm4p6Z
2018
Cited alongside, same era.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 2410–2419
2018
Cited alongside, same era.
K. Park and T. Mulc, “Css10: A collection of single speaker speech datasets for 10 languages,” Interspeech , 2019
2019
Later among the works it cites.
A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny, “Efficient lifelong learning with a-GEM,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=Hkf2_sC5FX
2019
Later among the works it cites.
M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y. Tu, , and G. Tesauro, “Learning to learn without forgetting by maximizing transfer and minimizing interference,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=B1gTShAct7
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
J. Xue, J. Han, T. Zheng, X. Gao, and J. Guo, “A multi-task learning framework for overcoming the catastrophic forgetting in automatic speech recognition,” 2019
2019
Cited alongside, same era.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie, C. Gulcehre, A. van den Oord, O. Vinyals, and N. de Freitas, “Sample efficient adaptive text-to-speech,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=rkzjUoAcFX
2019
Cited alongside, same era.
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019
2019
Cited alongside, same era.
2019
Later among the works it cites.
S. Sadhu and H. Hermansky, “Continual Learning in Automatic Speech Recognition,” in Proc. Interspeech 2020 , 2020, pp. 1246–1250. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2962
2020
Later among the works it cites.
T. Nekvinda and O. Dušek, “One Model, Many Languages: Meta-Learning for Multilingual Text-to-Speech,” in Proc. Interspeech 2020 , 2020, pp. 2972–2976. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2679
2020
Later among the works it cites.