Fetching the paper…
Reading the bibliography…
The present paper describes singing voice synthesis based on convolutional neural networks (CNNs).
“Mel log spectral approximation filter for speech synthesis,”
S. Imai, K. Sumita, and C. Furuichi, · 1983
Earlier work this paper cites.
“Static and dynamic error propagation networks with application to speech coding,”
A. J. Robinson and F. Fallside, · 1988
Earlier work this paper cites.
ITU-T Recommendation G.711
“Pulse code modulation (PCM) of voice frequencies,” · 1988
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Restructuring speech representations using the pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,”
H. Kawahara, M. K. Ikuyo, and A. d. Cheveigne, · 1999
Earlier work this paper cites.
“Speech parameter generation algorithms for HMM-based speech synthesis,”
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, · 2000
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
H. Zen, K. Tokuda, and A. W. Black, · 2009
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury, · 2012
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
H. Zen, A. Senior, and M. Schuster, · 2013
Earlier work this paper cites.
“A neural parametric singing synthesizer,”
M. Blaauw and J. Bonada, · 2013
Earlier work this paper cites.
“On the training aspects of deep neural network (DNN) for parametric TTS synthesis,”
Y. Qian, Y. Fan, W. Hu, and F. K. Soong, · 2014
Earlier work this paper cites.
“TTS synthesis with bidirectional LSTM based recurrent neural networks,”
Y. Fan, Y. Qian, F. Xie, and F. K. Soong, · 2014
Cited alongside, same era.
“Prosody contour prediction with long short-term memory, bidirectional, deep recurrent neural networks,”
R. Fernandez, A. Rendel, B. Ramabhadren, and R. Hoory, · 2014
Cited alongside, same era.
“Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis,”
H. Zen and H. Sak, · 2015
Cited alongside, same era.
“The effect of neural networks in statistical parametric speech synthesis,”
K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2015
Cited alongside, same era.
“Fully convolutional networks for semantic segmentation,”
J. Long, E. Shelhamer, and T. Darrell, · 2015
Cited alongside, same era.
“Singing voice synthesis based on deep neural networks,”
“Combining unidirectional long short-term memory with convolutional output layer for high-performance speech synthesis,”
W. Wang and B. Xu, · 2017
Later among the works it cites.
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. v. d. Oord, S. Dieleman, and K. Kavukcuoglu, · 2018
Later among the works it cites.
“FFTnet: A real-time speaker-dependent neural vocoder,”
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, · 2018
Later among the works it cites.
“WaveGlow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2018
Later among the works it cites.
“Recent development of the DNN-based singing voice synthesis system - Sinsy,”
Y. Hono, S. Murata, K. Nakamura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Nishimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2016
Cited alongside, same era.
“From HMMs to DNNs: where do the improvements come from?,”
O. Watts, G. E. Henter, T. Merritt, Z. Wu, and S. King, · 2016
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“SampleRNN: An unconditional end-to-end neural audio generation model,”
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, · 2016
Cited alongside, same era.
“Trajectory training considering global variance for speech synthesis based on neural networks,”
K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2016
Cited alongside, same era.
“Speaker-dependent WaveNet vocoder,”
A. Tamamori, T. Hayashi, K. Kovayashi, K. Takeda, and T. Toda, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, · 2018
Later among the works it cites.
“Close to human quality TTS with transformer,”
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, · 2018
Later among the works it cites.
“Mel-cepstrum-based quantization noise shaping applied to neural-network-based speech waveform synthesis,”
T. Yoshimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2018
Later among the works it cites.
“Reproducing high-quality singing voice with state-of-the-art AI technology,”
Techno-Speech,Inc., · 2018
Later among the works it cites.
“Singing voice synthesis using deep autoregressive neural networks for acoustic modeling,”
Y.-H. Yi, Y. Ai, Z.-H. Ling, and L.-R. Dai, · 2019
Closest in time.
“Adversarially trained end-to-end Korean singing voice synthesis system,”
J. Lee, H.-S. Choi, C.-B. Jeon, J. Koo, and K. Lee, · 2019
Closest in time.