Fetching the paper…
Reading the bibliography…
We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples.
“Speech-to-singing synthesis: Converting speaking voices to singing voices by controlling acoustic features unique to singing voices,”
Takeshi Saitou, Masataka Goto, Masashi Unoki, and Masato Akagi, · 2007
Earlier work this paper cites.
“Applying voice conversion to concatenative singing-voice synthesis,”
Fernando Villavicencio and Jordi Bonada, · 2010
Earlier work this paper cites.
“Statistical singing voice conversion with direct waveform modification based on the spectrum differential,”
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura, · 2014
Earlier work this paper cites.
“Statistical singing voice conversion based on direct waveform modification with global variance,”
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura, · 2015
Earlier work this paper cites.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Expressive singing synthesis based on unit selection for the singing synthesis challenge 2016.,”
Jordi Bonada, Martí Umbert, and Merlijn Blaauw, · 2016
Earlier work this paper cites.
“Singing voice synthesis based on deep neural networks.,”
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda, · 2016
Cited alongside, same era.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
“World: a vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Cited alongside, same era.
“A neural parametric singing synthesizer modeling timbre and expression from natural songs,”
Merlijn Blaauw and Jordi Bonada, · 2017
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
“Accurate, large minibatch sgd: Training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Later among the works it cites.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Later among the works it cites.
“Data efficient voice cloning for neural singing synthesis,”
Merlijn Blaauw, Jordi Bonada, and Ryunosuke Daido, · 2019
Closest in time.
“Unsupervised singing voice conversion,”
Eliya Nachmani and Lior Wolf, · 2019
Closest in time.
“Durian: Duration informed attention network for multimodal synthesis,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chengzhu Yu, Heng Lu, Na Hu, Meng Yu, Chao Weng, Kun Xu, Peng Liu, Deyi Tuo, Shiyin Kang, Guangzhi Lei, et al., · 2019
Closest in time.