Fetching the paper…
Reading the bibliography…
This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly.
“A learning algorithm for continually running fully recurrent neural networks,”
Ronald J Williams and David Zipser, · 1989
Earlier work this paper cites.
“Algorithms for non-negative matrix factorization,”
Daniel D Lee and H Sebastian Seung, · 2001
Earlier work this paper cites.
“The CMU arctic speech databases,”
J. Kominek and A. W. Black, · 2004
Earlier work this paper cites.
“A family of symmetric distributions on the circle,”
M. C. Jones and Arthur Pewsey, · 2005
Earlier work this paper cites.
“Supervised and semi-supervised separation of sounds from single-channel mixtures,”
P. Smaragdis, B. Raj, and M. Shashanka, · 2007
Earlier work this paper cites.
“INTERSPEECH 2014 special session: Phase importance in speech processing applications,”
Pejman Mowlaee, Rahim Saeidi, and Yannis Stylianou, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Wavenet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
“WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Cited alongside, same era.
“Direct modeling of frequency spectra and waveform generation based on phase recovery for DNN-based speech synthesis,”
Shinji Takaki, Hirokazu Kameoka, and Junichi Yamagishi, · 2017
Cited alongside, same era.
“Speaker-dependent WaveNet vocoder,”
Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, and Tomoki Toda, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Closest in time.
“A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,”
Xin Wang, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela, and Junichi Yamagishi, · 2018
Closest in time.
“Clarinet: Parallel wave generation in end-to-end text-to-speech,”
Wei Ping, Kainan Peng, and Jitong Chen, · 2018
Closest in time.
“Can we steal your vocal identity from the internet?: Initial investigation of cloning obama’s voice using gan, wavenet and low-quality found data,”
Jaime Lorenzo-Trueba, Fuming Fang, Xin Wang, Isao Echizen, Junichi Yamagishi, and Tomi Kinnunen, · 2018
Closest in time.
“Autoregressive neural f0 model for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al., · 2017
Cited alongside, same era.
Closest in time.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2019
Closest in time.