Fetching the paper…
Reading the bibliography…
Recently, end-to-end Korean singing voice systems have been designed to generate realistic singing voices.
W. Frank, “Individual comparisons by ranking methods,” Biometrics Bulletin , vol. 1, no. 6, pp. 80–83, 1945
1945
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proceedings NIPS Neural Information Processing Systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
H. Y. Gu and J. K. He, “Singing-voice synthesis using demi-syllable unit selection,” in Proceedings ICMLC 2016 International Conference on Machine Learnign and Cybernetics , 2016, pp. 654–659
2016
Earlier work this paper cites.
M. Nishimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Singing voice synthesis based on deep neural networks,” in Proceedings INTERSPEECH 2016 International Speech Communication Association , 2016, pp. 2478–2482
2016
Earlier work this paper cites.
D. A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (elus),” in Proceedings ICLR International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE TRANSACTIONS on Information and Systems , vol. 99, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proceedings INTERSPEECH 2017 International Speech Communication Association , 2017, p. 4006–4010
2017
Earlier work this paper cites.
M. Blaauw and J. Bonada, “A neural parametric singing synthesizer modeling timbre and expression from natural songs,” Applied Sciences , vol. 7, no. 12, p. 1333, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings NIPS Neural Information Processing Systems 30 , 2017, pp. 6000–6010
2017
Cited alongside, same era.
J. Kim, H. Choi, J. Park, M. Hahn, S.-J. Kim, and J.-J. Kim, “Korean singing voice synthesis based on an lstm recurrent neural network.” in Proceedings INTERSPEECH 2018 International Speech Communication Association , 2018, p. 1551–1555
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, and R. S.-R. et al, “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in Proceedings ICASSP 2018 IEEE International Conference on Acoustics, Speech and Signal Processing , 2018, pp. 4779–4783
S. Choi, W. Kim, S. Park, S. Yong, and J. Nam, “Korean singing voice synthesis based on auto-regressive boundary equilibrium gan,” in Proceedings ICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal Processing , 2020, pp. 7234–7238
2020
Later among the works it cites.
M. Blaauw and J. Bonada, “Sequence-to-sequence singing synthesis using the feed-forward transformer,” in Proceedings ICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal Processing , 2020, pp. 7229–7233
2020
Later among the works it cites.
P. Lu, J. Wu, J. Luan, X. Tan, and L. Zhou, “Xiaoicesing: A high-quality and integrated singing voice synthesis system,” in Proceedings INTERSPEECH 2020 International Speech Communication Association , 2020, pp. 1306–1310
2020
Later among the works it cites.
J. Wu and J. Luan, “Adversarially trained multi-singer sequence-to-sequence singing synthesizer,” in Proceedings INTERSPEECH 2020 International Speech Communication Association , 2020, pp. 1296–1300
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,” in Proceedings ICASSP 2018 IEEE International Conference on Acoustics, Speech and Signal Processing , 2018, pp. 4784–4788
2018
Cited alongside, same era.
T. Miyato and M. Koyama, “cgans with projection discriminator,” in Proceedings ICLR 2018 International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
J. Lee, H. S. Choi, C. B. Jeon, J. Koo, and K. Lee, “Adversarially trained end-to-end korean singing voice synthesis system,” in Proceedings INTERSPEECH 2019 International Speech Communication Association , 2019, pp. 803–806
2019
Cited alongside, same era.
Y. Ren, Y. Ruan., X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in Proceedings NIPS Neural Information Processing Systems 32 , 2019, pp. 3165–3174
2019
Cited alongside, same era.
I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le, “Attention augmented convolutional networks,” in Proceedings ICCV 2019 IEEE International Conference on Computer Vision , 2019, pp. 3286–3295
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. M. Association, “Midi manufacturers association,” in https://www.midi.org
Cited in the paper.
Later among the works it cites.
A. Gulati, J. Qin, C. Chiu, and e. a. N. Parmar, Y. Zhang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proceedings INTERSPEECH 2020 International Speech Communication Association , 2020, pp. 5036–5040
2020
Later among the works it cites.
2021
Closest in time.
Y. Gu, X. Yin, Y. Rao, Y. Wan, B. Tang, and Y. Zhang, “Bytesing: A chinese singing voice synthesis system using duration allocated encoder-decoder acoustic models and wavernn vocoders,” in Proceedings ISCSLP 12th International Symposium on Chinese Spoken Language Processing , 2021, pp. 1–5
2021
Closest in time.
R. Yamamoto, E. Song, and J. M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proceedings ICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal Processing , 2021, pp. 6199–6203
2021
Closest in time.
R. Yamamoto, E. Song, M. J. Hwang, and J. M. Kim, “Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discriminators,” in Proceedings ICASSP 2021 IEEE International Conference on Acoustics, Speech and Signal Processing , 2021, pp. 6039–6043
2021
Closest in time.