Fetching the paper…
Reading the bibliography…
Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure.
“Variable frequency electric circuit theory with application to the theory of frequency-modulation,”
John R Carson and Thornton C Fry, · 1937
Earlier work this paper cites.
“A mixed-source model for speech compression and synthesis,”
John Makhoul, R Viswanathan, Richard Schwartz, and AWF Huggins, · 1978
Earlier work this paper cites.
“A tone oriented voice excited vocoder,”
Per Hedelin, · 1981
Earlier work this paper cites.
“A four-parameter model of glottal flow,”
Gunnar Fant, Johan Liljencrants, and Qi-guang Lin, · 1985
Earlier work this paper cites.
“Speech analysis/synthesis based on a sinusoidal representation,”
Robert McAulay and Thomas Quatieri, · 1986
Earlier work this paper cites.
“Multiband excitation vocoder,”
D. W. Griffin and J. S. Lim, · 1988
Earlier work this paper cites.
“Mel-generalized cepstral analysis a unified approach,”
Keiichi Tokuda, Takao Kobayashi, Takashi Masuko, and Satoshi Imai, · 1994
Earlier work this paper cites.
“Algorithms for non-negative matrix factorization,”
Daniel D Lee and H Sebastian Seung, · 2001
Earlier work this paper cites.
“XIMERA: A new TTS from ATR based on corpus-based technologies,”
Hisashi Kawai, Tomoki Toda, Jinfu Ni, Minoru Tsuzaki, and Keiichi Tokuda, · 2004
Earlier work this paper cites.
Supervised Sequence Labelling with Recurrent Neural Networks
Alex Graves, · 2008
Cited alongside, same era.
“Introducing CURRENT: The Munich open-source CUDA recurrent neural network toolkit,”
Felix Weninger, Johannes Bergmann, and Björn Schuller, · 2015
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
“Improved variational inference with inverse autoregressive flow,”
Diederik P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling, · 2016
Cited alongside, same era.
“WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Cited alongside, same era.
“Direct modeling of frequency spectra and waveform generation based on phase recovery for DNN-based speech synthesis,”
Shinji Takaki, Hirokazu Kameoka, and Junichi Yamagishi, · 2017
Later among the works it cites.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Closest in time.
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Dan Belov, and Demis Hassabis, · 2018
Closest in time.
“A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,”
Xin Wang, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela, and Junichi Yamagishi, · 2018
Closest in time.
“Investigation of WaveNet for text-to-speech synthesis,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network,”
Lauri Juvela, Bajibabu Bollepalli, Manu Airaksinen, and Paavo Alku, · 2016
Cited alongside, same era.
“Speaker-dependent WaveNet vocoder,”
Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, and Tomoki Toda, · 2017
Cited alongside, same era.
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2018
Closest in time.
“Investigating accuracy of pitch-accent annotations in neural-network-based speech synthesis and denoising effects,”
Hieu-Thi Luong, Xin Wang, Junichi Yamagishi, and Nobuyuki Nishizawa, · 2018
Closest in time.
“Clarinet: Parallel wave generation in end-to-end text-to-speech,”
Wei Ping, Kainan Peng, and Jitong Chen, · 2019
Closest in time.
“STFT spectral loss for training a neural speech waveform model,”
Shinji Takaki, Toru Nakashika, Xin Wang, and Junichi Yamagishi, · 2019
Closest in time.