Fetching the paper…
Reading the bibliography…
This paper presents a method that generates expressive singing voice of Peking opera.
“Speech synthesis,”
C Lochbaum and J Kelly, · 1962
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“pyin: A fundamental frequency estimator using probabilistic threshold distributions,”
Matthias Mauch and Simon Dixon, · 2014
Earlier work this paper cites.
“Automatic identification of emotional cues in chinese opera singing,”
Dawn AA Black, Ma Li, and Mi Tian, · 2014
Earlier work this paper cites.
“Pitch contour segmentation for computer-aided jinju singing training,”
Rong Gong, Yile Yang, and Xavier Serra, · 2016
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“A neural parametric singing synthesizer modeling timbre and expression from natural songs,”
Merlijn Blaauw and Jordi Bonada, · 2017
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“Recent development of the DNN-based singing voice synthesis system — sinsy,”
Yukiya Hono, Shumma Murata, Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda, · 2018
Cited alongside, same era.
“Sequential generation of singing f0 contours from musical note sequences based on wavenet,”
Yusuke Wada, Ryo Nishikimi, Eita Nakamura, Katsutoshi Itoyama, and Kazuyoshi Yoshii, · 2018
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Deep voice 3: 2000-speaker neural text-to-speech,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2018
Cited alongside, same era.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
“Wgansing: A multi-voice singing voice synthesizer based on the wasserstein-gan,”
Pritish Chandna, Merlijn Blaauw, Jordi Bonada, and Emilia Gomez, · 2019
Closest in time.
“Melnet: A generative model for audio in the frequency domain,”
Sean Vasquez and Mike Lewis, · 2019
Closest in time.
“Jingju a cappella recordings collection,” 2019
Rong Gong, Rafael Caro, and Tiange Zhu, · 2019
Closest in time.
“Jingju a cappella singing dataset part1,” https://doi.org/10.5281/zenodo.1323561
Rong Gong, Rafael Caro Repetto, Yile Yang, and Xavier Serra, · 2019
Closest in time.
“Jingju a cappella singing dataset part2,” https://doi.org/10.5281/zenodo.1421692
Rong Gong, Rafael Caro Repetto, and Xavier Serra, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Durian: Duration informed attention network for multimodal synthesis,”
Chengzhu Yu, Heng Lu, Na Hu, Meng Yu, Chao Weng, Kun Xu, Peng Liu, Deyi Tuo, Shiyin Kang, Guangzhi Lei, et al., · 2019
Cited alongside, same era.
“Singing voice synthesis using deep autoregressive neural networks for acoustic modeling,”
Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, and Li-Rong Dai, · 2019
Cited alongside, same era.
“Enhanced virtual singers generation by incorporating singing dynamics to personalized text-to-speech-to-singing,”
Kantapon Kaewtip, Fernando Villavicencio, Fang-Yu Kuo, Mark Harvilla, Iris Ouyang, and Pierre Lanchantin, · 2019
Cited alongside, same era.
“Jingju a cappella singing dataset part3,” https://doi.org/10.5281/zenodo.1286350
Rong Gong and Xavier Serra, · 2019
Closest in time.
“Speech signal processing toolkit (sptk),” http://sp-tk.sourceforge.net/
K Tokuda, K Oura, A Tamamori, S Sako, H Zen, T Nose, T Takahashi, J Yamagishi, and Y Nankaku, · 2019
Closest in time.