Fetching the paper…
Reading the bibliography…
Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically mel-spectrograms), followed by a 'neural vocoder' model which produces the time-domain speech waveform from this intermediate speech representation.
“Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system STRAIGHT,”
Hideki Kawahara, Jo Estill, and Osamu Fujimura, · 2001
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
Heiga Zen, Keiichi Tokuda, and Alan W Black, · 2009
Earlier work this paper cites.
“A unified trajectory tiling approach to high quality speech rendering,”
Yao Qian, Frank K Soong, and Zhi-Jie Yan, · 2012
Earlier work this paper cites.
Nishant Prateek, Mateusz Lajszczak, Roberto Barra-Chicote, Thomas Drugman, Jaime Lorenzo-Trueba, Thomas Merritt, Srikanth Ronanki, and Trevor Wood, · 2012
Earlier work this paper cites.
“Auto-encoding variational bayes,”
Diederik P. Kingma and Max Welling, · 2014
Earlier work this paper cites.
“Acoustic modeling in statistical parametric speech synthesis-from HMM to LSTM-RNN,”
Heiga Zen, · 2015
Earlier work this paper cites.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Earlier work this paper cites.
“Deep neural network-guided unit selection synthesis,”
Thomas Merritt, Robert AJ Clark, Zhizheng Wu, Junichi Yamagishi, and Simon King, · 2016
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
“Improved variational inference with inverse autoregressive flow,”
Diederik P. Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling, · 2016
Cited alongside, same era.
Overcoming the limitations of statistical parametric speech synthesis
Thomas Merritt, · 2017
Cited alongside, same era.
“Google’s Next-Generation Real-Time Unit-Selection Synthesizer Using Sequence-to-Sequence LSTM-Based Autoencoders,”
Vincent Wan, Yannis Agiomyrgiannakis, Hanna Silen, and Jakub Vit, · 2017
Cited alongside, same era.
“Deep voice: Real-time neural text-to-speech,”
Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi, · 2017
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
Aaron Van Den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Van Den Driessche, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Dan Belov, and Demis Hassabis, · 2018
Later among the works it cites.
“Towards achieving robust universal neural vocoding,”
Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, and Vatsal Aggarwal, · 2019
Later among the works it cites.
“Clarinet: Parallel wave generation in end-to-end text-to-speech,”
Wei Ping, Kainan Peng, and Jitong Chen, · 2019
Later among the works it cites.
“Waveglow: A Flow-based Generative Network for Speech Synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Later among the works it cites.
“Adversarial audio synthesis,”
Chris Donahue, Julian McAuley, and Miller Puckette, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimber, Aäron Van Den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Cited alongside, same era.
“Learning latent representations for style control and transfer in end-to-end speech synthesis,”
Ya-Jie Zhang, Shifeng Pan, Lei He, and Zhen-Hua Ling, · 2019
Later among the works it cites.
“Using generative modelling to produce varied intonation for speech synthesis,”
Zack Hodari, Oliver Watts, and Simon King, · 2019
Later among the works it cites.