Fetching the paper…
Reading the bibliography…
In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites.
Signal estimation from modified short-time Fourier transform
Daniel Griffin and Jae Lim. 1984 · 1984
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database. In ICASSP 1996 , Vol. 1. IEEE, 373–376
Andrew J Hunt and Alan W Black. 1996 · 1996
Earlier work this paper cites.
Praat, a system for doing phonetics by computer
Paul Boersma et al · 2002
Earlier work this paper cites.
Tutorial on training recurrent neural networks, covering BPPT, RTRL, EKF and the" echo state network" approach . Vol. 5
Herbert Jaeger. 2002 · 2002
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine learning . 129–136
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007 · 2007
Earlier work this paper cites.
Alignment of lyrics with accompanied singing audio based on acoustic-phonetic vowel likelihood modeling
Yu-Ren Chien, Hsin-Min Wang, Shyh-Kang Jeng, Yu-Ren Chien, Hsin-Min Wang, and Shyh-Kang Jeng. 2016 · 2008
Earlier work this paper cites.
Automatic alignment of music audio and lyrics. In Proceedings of the 11th Int. Conference on Digital Audio Effects (DAFx-08)
Annamaria Mesaros and Tuomas Virtanen. 2008 · 2008
Earlier work this paper cites.
Clueweb09 data set
Jamie Callan, Mark Hoy, Changkuk Yoo, and Le Zhao. 2009 · 2009
Earlier work this paper cites.
LETOR: A benchmark collection for research on learning to rank for information retrieval
Tao Qin, Tie-Yan Liu, Jun Xu, and Hang Li. 2010 · 2010
Earlier work this paper cites.
LyricSynchronizer: Automatic synchronization system between musical audio signals and lyrics
Hiromasa Fujihara, Masataka Goto, Jun Ogata, and Hiroshi G Okuno. 2011 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
THCHS-30 : A Free Chinese Speech Corpus
Zhiyong Zhang Dong Wang, Xuewei Zhang. 2015 · 2015
Earlier work this paper cites.
Expression control in singing voice synthesis: Features, approaches, evaluation, and challenges
Marti Umbert, Jordi Bonada, Masataka Goto, Tomoyasu Nakano, and Johan Sundberg. 2015 · 2015
Earlier work this paper cites.
Singing Voice Synthesis Based on Deep Neural Networks.. In Interspeech . 2478–2482
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Merlin: An Open Source Neural Network Speech Synthesis System.. In SSW . 202–207
Zhizheng Wu, Oliver Watts, and Simon King. 2016 · 2016
Cited alongside, same era.
A neural parametric singing synthesizer modeling timbre and expression from natural songs
Merlijn Blaauw and Jordi Bonada. 2017 · 2017
Cited alongside, same era.
Knowledge-based probabilistic modeling for tracking lyrics in music audio signals
Georgi Dzhambazov et al · 2017
Cited alongside, same era.
Webvision database: Visual learning and understanding from web data
Wen Li, Limin Wang, Wei Li, Eirikur Agustsson, and Luc Van Gool. 2017 · 2017
Cited alongside, same era.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi. In Interspeech . 498–502
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Cited alongside, same era.
Spleeter: A fast and state-of-the art music source separation tool with pre-trained models. In Proc. International Society for Music Information Retrieval Conference
Romain Hennequin, Anis Khlif, Felix Voituret, and Manuel Moussalam. 2019 · 2019
Later among the works it cites.
Singing Voice Synthesis Based on Generative Adversarial Networks. In ICASSP 2019 . IEEE, 6955–6959
Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda. 2019 · 2019
Later among the works it cites.
Adversarially Trained End-to-end Korean Singing Voice Synthesis System
Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, and Kyogu Lee. 2019 · 2019
Later among the works it cites.
Singing voice synthesis based on convolutional neural networks
Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need. In Advances in Neural Information Processing Systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Semi-supervised Lyrics and Solo-singing Alignment.. In ISMIR . 600–607
Chitralekha Gupta, Rong Tong, Haizhou Li, and Ye Wang. 2018 · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018 · 2018
Cited alongside, same era.
Korean Singing Voice Synthesis System based on an LSTM Recurrent Neural Network. In INTERSPEECH 2018 . ISCA
Juntae Kim, Heejin Choi, Jinuk Park, Sangjin Kim, Jongjin Kim, and Minsoo Hahn. 2018 · 2018
Cited alongside, same era.
EMPHASIS: An emotional phoneme-based acoustic model for speech synthesis system
Hao Li, Yongguo Kang, and Zhenyu Wang. 2018 · 2018
Cited alongside, same era.
A Study on Correlates of Acoustic Features to Emotional Singing Voice Synthesis
Thi Hao Nguyen. 2018 · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In ICASSP 2018 . IEEE, 4779–4783
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
Wei Ping, Kainan Peng, and Jitong Chen. 2019 · 2019
Later among the works it cites.
FastSpeech: Fast, Robust and Controllable Text to Speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Automatic Lyrics-to-audio Alignment on Polyphonic Music Using Singing-adapted Acoustic Models. In ICASSP 2019 . IEEE, 396–400
Bidisha Sharma, Chitralekha Gupta, Haizhou Li, and Ye Wang. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in neural information processing systems . 5754–5764
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling
Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, and Li-Rong Dai. 2019 · 2019
Later among the works it cites.
LibriTTS: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Later among the works it cites.
Liqiang Zhang, Chengzhu Yu, Heng Lu, Chao Weng, Yusong Wu, Xiang Xie, Zijin Li, and Dong Yu. 2019 · 2019
Later among the works it cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020 · 2020
Closest in time.
XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System
Peiling Lu, Jie Wu, Jian Luan, Xu Tan, and Li Zhou. 2020 · 2020
Closest in time.
FastSpeech 2: Fast and High-Quality End-to-End Text-to-Speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Closest in time.