Fetching the paper…
Reading the bibliography…
Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel.
“A maximization technique occurring in the statistical analysis of probabilistic functions of markov chains,”
Leonard E Baum, Ted Petrie, George Soules, and Norman Weiss, · 1970
Earlier work this paper cites.
Mixture density networks
Christopher M Bishop, · 1994
Earlier work this paper cites.
Text-to-speech synthesis
Paul Taylor, · 2009
Earlier work this paper cites.
“Speech synthesis based on hidden markov models,”
Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, and Keiichiro Oura, · 2013
Earlier work this paper cites.
“Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis,”
Heiga Zen and Andrew Senior, · 2014
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Char2wav: End-to-end speech synthesis,”
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio, · 2017
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Cited alongside, same era.
“Deep voice: Real-time neural text-to-speech,”
Sercan Ö Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al., · 2017
Cited alongside, same era.
“Deep voice 2: Multi-speaker neural text-to-speech,”
Sercan Ö Arik, Gregory Diamos, Andrew Gibiansky, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“The lj speech dataset,” https://keithito.com/ LJ-Speech-Dataset/, 2017
Keith Ito, · 2017
Cited alongside, same era.
“Voiceloop: Voice fitting and synthesis via a phonological loop,”
“Deep voice 3: 2000-speaker neural text-to-speech,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2018
Later among the works it cites.
“Clarinet: Parallel wave generation in end-to-end text-to-speech,”
Wei Ping, Kainan Peng, and Jitong Chen, · 2018
Later among the works it cites.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Later among the works it cites.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, Ming Liu, and M Zhou, · 2019
Later among the works it cites.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani, · 2018
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao, · 2019
Later among the works it cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Later among the works it cites.