Fetching the paper…
Reading the bibliography…
This paper presents ByteSing, a Chinese singing voice synthesis (SVS) system based on duration allocated Tacotron-like acoustic models and WaveRNN neural vocoders.
K. Saino, H. Zen, Y. Nankaku, A. Lee, and K. Tokuda, “An HMM-based singing voice synthesis system,” in
2006
Earlier work this paper cites.
M. Good, “MusicXML in commercial applications,” pp. 9–20, 2006
2006
Earlier work this paper cites.
K. Oura, A. Mase, T. Yamada, S. Muto, Y. Nankaku, and K. Tokuda, “Recent development of the HMM-based singing voice synthesis system—Sinsy,” in
2010
Earlier work this paper cites.
A. Graves, “Generating sequences with recurrent neural networks,”
2013
Earlier work this paper cites.
Z.-H. Ling, S.-Y. Kang, H. Zen, A. Senior, M. Schuster, X.-J. Qian, H. Meng, and L. Deng, “Deep Learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends,”
2015
Earlier work this paper cites.
M. Nishimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Singing voice synthesis based on deep neural networks.” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in
2017
Cited alongside, same era.
Y. Hono, S. Murata, K. Nakamura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Recent development of the DNN-based singing voice synthesis system — Sinsy,” in
2018
Cited alongside, same era.
J. Kim, H. Choi, J. Park, M. Hahn, S.-J. Kim, and J.-J. Kim, “Korean singing voice synthesis based on an LSTM recurrent neural network.” in
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
2019
Later among the works it cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in
2019
Later among the works it cites.
C. Yu, H. Lu, N. Hu, M. Yu, C. Weng, K. Xu, P. Liu, D. Tuo, S. Kang, G. Lei
2019
Later among the works it cites.
2019
Later among the works it cites.
L. Zhang, C. Yu, H. Lu, C. Weng, Y. Wu, X. Xie, Z. Li, and D. Yu, “Learning singing from speech,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Later among the works it cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in
2019
Later among the works it cites.
2019
Later among the works it cites.