Fetching the paper…
Reading the bibliography…
We propose a neural text-to-speech (TTS) model that can imitate a new speaker's voice using only a small amount of speech sample.
Signal estimation from modified short-time fourier transform
Griffin, Daniel and Lim, Jae · 1984
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, David E, Hinton, Geoffrey E, and Williams, Ronald J · 1985
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Cited alongside, same era.
Wavenet: A generative model for raw audio
van den Oord, Aaron, Dieleman, Sander, Zen, Heiga, Simonyan, Karen, Vinyals, Oriol, Graves, Alexander, Kalchbrenner, Nal, Senior, Andrew, and Kavukcuoglu, Koray · 2016
Cited alongside, same era.
Emotional end-to-end neural speech synthesizer
Lee, Younggun, Rabiee, Azam, and Lee, Soo-Young · 2017
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
Mehri, Soroush, Kumar, Kundan, Gulrajani, Ishaan, Kumar, Rithesh, Jain, Shubham, Sotelo, Jose, Courville, Aaron, and Bengio, Yoshua · 2017
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
Arık, Sercan Ö., Chrzanowski, Mike, Coates, Adam, Diamos, Gregory, Gibiansky, Andrew, Kang, Yongguo, Li, Xian, Miller, John, Ng, Andrew, Raiman, Jonathan, Sengupta, Shubho, and Shoeybi, Mohammad
Cited in the paper.
Deep voice 2: Multi-speaker neural text-to-speech
Arık, Sercan O, Diamos, Gregory, Gibiansky, Andrew, Miller, John, Peng, Kainan, Ping, Wei, Raiman, Jonathan, and Zhou, Yanqi
Cited in the paper.
Webrtc voice activity detector
Cited in the paper.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Shen, Jonathan, Pang, Ruoming, Weiss, Ron J, Schuster, Mike, Jaitly, Navdeep, Yang, Zongheng, Chen, Zhifeng, Zhang, Yu, Wang, Yuxuan, Skerry-Ryan, RJ, et al · 2017
Later among the works it cites.
Char2wav: End-to-end speech synthesis
Sotelo, Jose, Mehri, Soroush, Kumar, Kundan, Santos, Joao Felipe, Kastner, Kyle, Courville, Aaron, and Bengio, Yoshua · 2017
Later among the works it cites.
Tacotron: Towards end-to-end speech synthesis
Wang, Yuxuan, Skerry-Ryan, R.J., Stanton, Daisy, Wu, Yonghui, Weiss, Ron J., Jaitly, Navdeep, Yang, Zongheng, Xiao, Ying, Chen, Zhifeng, Bengio, Samy, Le, Quoc, Agiomyrgiannakis, Yannis, Clark, Rob, and Saurous, Rif A · 2017
Later among the works it cites.
Deep voice 3: 2000-speaker neural text-to-speech
Ping, Wei, Peng, Kainan, Gibiansky, Andrew, Arik, Sercan O., Kannan, Ajay, Narang, Sharan, Raiman, Jonathan, and Miller, John · 2018
Closest in time.
Voiceloop: Voice fitting and synthesis via a phonological loop
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taigman, Yaniv, Wolf, Lior, Polyak, Adam, and Nachmani, Eliya · 2018
Closest in time.