Fetching the paper…
Reading the bibliography…
With rapid progress in neural text-to-speech (TTS) models, personalized speech generation is now in high demand for many applications.
Learning to Learn
Thrun, S. and Pratt, L. Y. (eds.) · 1998
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, L. V. D. and Hinton, G. E · 2008
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2014
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
McFee, B., Raffel, C., Liang, D., Ellis, D., McVicar, M., Battenberg, E., and Nieto, O · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Deep speech 2 : End-to-end speech recognition in english and mandarin
Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Chen, J., Chrzanowski, M., Coates, A., Diamos, G., Elsen, E., Engel, J. H., Fan, L., Fougner, C., Hannun, A. Y., Jun, B., Han, T., LeGresley, P., Li, X., Lin, L., Narang, S., Ng, A. Y., Ozair, S., Prenger, R., Qian, S., Raiman, J., Satheesh, S., Seetapun, D., Sengupta, S., Wang, C., Wang, Y., Wang, Z., Xiao, B., Xie, Y., Yogatama, D., Zhan, J., and Zhu, Z · 2016
Earlier work this paper cites.
Ba, J., Kiros, J., and Hinton, G. E · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
One-shot generalization in deep generative models
Rezende, D. J., Mohamed, S., Danihelka, I., Gregor, K., and Wierstra, D · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Arik, S. Ö., Chrzanowski, M., Coates, A., Diamos, G. F., Gibiansky, A., Kang, Y., Li, X., Miller, J., Ng, A. Y., Raiman, J., Sengupta, S., and Shoeybi, M · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Deep voice 2: Multi-speaker neural text-to-speech
Gibiansky, A., Arik, S. Ö., Diamos, G. F., Miller, J., Peng, K., Ping, W., Raiman, J., and Zhou, Y · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Mao, X., Li, Q., Xie, H., Lau, R. Y. K., Wang, Z., and Smolley, S. P · 2017
Earlier work this paper cites.
Deep voice 3: 2000-speaker neural text-to-speech
Ping, W., Peng, K., Gibiansky, A., Arik, S. Ö., Kannan, A., Narang, S., Raiman, J., and Miller, J · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R. S · 2017
Earlier work this paper cites.
Char2wav: End-to-end speech synthesis
Sotelo, J., Mehri, S., Kumar, K., Santos, J. F., Kastner, K., Courville, A. C., and Bengio, Y · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q. V., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A · 2017
Cited alongside, same era.
Neural voice cloning with a few samples
Arik, S. Ö., Chen, J., Peng, K., Ping, W., and Zhou, Y · 2018
Cited alongside, same era.
Few-shot generative modelling with generative matching networks
Bartunov, S. and Vetrov, D. P · 2018
Cited alongside, same era.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Lopez-Moreno, I., and Wu, Y · 2018
Sample efficient adaptive text-to-speech
Chen, Y., Assael, Y. M., Shillingford, B., Budden, D., Reed, S. E., Zen, H., Wang, Q., Cobo, L. C., Trask, A., Laurie, B., Gülçehre, Ç., van den Oord, A., Vinyals, O., and de Freitas, N · 2019
Later among the works it cites.
Figr: Few-shot image generation with reptile
Clouâtre, L. and Demers, M · 2019
Later among the works it cites.
Hierarchical generative modeling for controllable speech synthesis
Hsu, W., Zhang, Y., Weiss, R. J., Zen, H., Wu, Y., Wang, Y., Cao, Y., Jia, Y., Chen, Z., Shen, J., Nguyen, P., and Pang, R · 2019
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Later among the works it cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brébisson, A., Bengio, Y., and Courville, A. C · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
cgans with projection discriminator
Miyato, T. and Koyama, M · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Fitting new speakers based on a short untranscribed sample
Nachmani, E., Polyak, A., Taigman, Y., and Wolf, L · 2018
Cited alongside, same era.
TADAM: task dependent adaptive metric for improved few-shot learning
Oreshkin, B. N., López, P. R., and Lacoste, A · 2018
Cited alongside, same era.
Few-shot autoregressive density estimation: Towards learning to learn distributions
Reed, S. E., Chen, Y., Paine, T., van den Oord, A., Eslami, S. M. A., Rezende, D. J., Vinyals, O., and de Freitas, N · 2018
Cited alongside, same era.
Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Ryan, R., Saurous, R. A., Agiomyrgiannakis, Y., and Wu, Y · 2018
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T · 2019
Later among the works it cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
Yamagishi, J., Veaux, C., and Macdonald, K · 2019
Later among the works it cites.
Libritts: A corpus derived from librispeech for text-to-speech
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Later among the works it cites.
Multispeech: Multi-speaker text to speech with transformer
Chen, M., Tan, X., Ren, Y., Xu, J., Sun, H., Zhao, S., and Qin, T · 2020
Later among the works it cites.
Fastpitch: Parallel text-to-speech with pitch prediction
La’ncucki, A · 2020
Later among the works it cites.
Mish: A self regularized non-monotonic activation function
Misra, D · 2020
Later among the works it cites.
BOFFIN TTS: few-shot speaker adaptation by bayesian optimization
Moss, H. B., Aggarwal, V., Prateek, N., González, J., and Barra-Chicote, R · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2020
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T · 2020
Later among the works it cites.
Adadurian: Few-shot adaptation for neural text-to-speech with durian
Zhang, Z., Tian, Q., Lu, H., Chen, L., and Liu, S · 2020
Later among the works it cites.
Adaspeech: Adaptive text to speech for custom voice
Chen, M., Tan, X., Li, B., Liu, Y., Qin, T., sheng zhao, and Liu, T.-Y · 2021
Closest in time.