Fetching the paper…
Reading the bibliography…
We propose a sequence-to-sequence singing synthesizer, which avoids the need for training data with pre-aligned phonetic and acoustic features.
“Generating sequences with recurrent neural networks,”
Alex Graves, · 2013
Earlier work this paper cites.
“Sequence level training with recurrent neural networks,”
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojchiech Zaremba, · 2016
Earlier work this paper cites.
“WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton, · 2016
Earlier work this paper cites.
“Tacotron: A fully end-to-end text-to-speech synthesis model,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc V. Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous, · 2017
Earlier work this paper cites.
“Convolutional sequence to sequence learning,”
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin, · 2017
Earlier work this paper cites.
“Speaker-dependent WaveNet vocoder,”
Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, and Tomoki Toda, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Language modeling with gated convolutional networks,”
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier, · 2017
Earlier work this paper cites.
“A neural parametric singing synthesizer modeling timbre and expression from natural songs,”
Merlijn Blaauw and Jordi Bonada, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, Rif A. Saurous, Yannis Agiomygiannakis, and Yonghui Wu, · 2018
Cited alongside, same era.
“Deep Voice 3: Scaling text-to-speech with convolutional sequence learning,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan Ö. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2018
Cited alongside, same era.
“Self-attentional acoustic models,”
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel, · 2018
Cited alongside, same era.
“Efficient trainable text-to-speech system based on deep convolutional networks with guided attention,”
Hideyuki Tachibana, Katsuya Uenoyama, and Shunsuke Aihara, · 2018
Cited alongside, same era.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2019
Closest in time.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Closest in time.
“Adversarially trained end-to-end Korean singing voice synthesis system,”
Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, and Kyogu Lee, · 2019
Closest in time.
“Analysing deep learning-spectral envelope prediction methods for singing synthesis,”
Frederik Bous and Axel Roebel, · 2019
Closest in time.
“Singing voice synthesis using deep autoregressive neural networks for acoustic modeling,”
Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, and Li-Rong Dai, · 2019
Closest in time.
“Singing voice synthesis based on convolutional neural networks,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao, · 2019
Cited alongside, same era.
“FastSpeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Cited alongside, same era.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Cited alongside, same era.
Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda, · 2019
Closest in time.
“Singing voice synthesis based on generative adversarial networks,”
Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda, · 2019
Closest in time.
“WGANSing: A multi-voice singing voice synthesizer based on the Wasserstein-GAN,”
Pritish Chandna, Merlijn Blaauw, Jordi Bonada, and Emilia Gómez, · 2019
Closest in time.