Fetching the paper…
Reading the bibliography…
Recent developments in deep learning have significantly improved the quality of synthesized singing voice audio.
“Sample-based singing voice synthesizer using spectral models and source-filter decomposition,”
J. Bonada, Alex Loscos, Oscar Mayor, and H. Kenmochi, · 2003
Earlier work this paper cites.
“An hmm-based singing voice synthesis system,”
Keijiro Saino, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, and Keiichi Tokuda, · 2006
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“World: A vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Earlier work this paper cites.
“Layer normalization,” 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton, · 2016
Earlier work this paper cites.
“A neural parametric singing synthesizer modeling timbre and expression from natural songs,”
Merlijn Blaauw and Jordi Bonada, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Korean grapheme-to-phoneme analyzer (kog2p),” https://github.com/scarletcho/KoG2P , 2017
Yejin Cho, · 2017
Cited alongside, same era.
“Wgansing: A multi-voice singing voice synthesizer based on the wasserstein-gan,”
Pritish Chandna, Merlijn Blaauw, Jordi Bonada, and Emilia Gómez, · 2019
Cited alongside, same era.
“Adversarially trained end-to-end korean singing voice synthesis system,”
Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, and Kyogu Lee, · 2019
Cited alongside, same era.
“Generalization in generation: A closer look at exposure bias,”
Florian Schmidt, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Reformer: The efficient transformer,”
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya, · 2020
Later among the works it cites.
“Gaussian error linear units (gelus),” 2020
Dan Hendrycks and Kevin Gimpel, · 2020
Later among the works it cites.
“Children’s song dataset for singing voice research,”
Soonbeom Choi, Wonil Kim, Saebyul Park, Sangeon Yong, and Juhan Nam, · 2020
Later among the works it cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Later among the works it cites.
“Mlp-mixer: An all-mlp architecture for vision,” 2021
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy, · 2021
Closest in time.
“Pay attention to mlps,” 2021
Hanxiao Liu, Zihang Dai, David R. So, and Quoc V. Le, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Sequence-to-sequence singing synthesis using the feed-forward transformer,”
Merlijn Blaauw and Jordi Bonada, · 2020
Cited alongside, same era.
“Korean singing voice synthesis based on auto-regressive boundary equilibrium gan,”
Soonbeom Choi, Wonil Kim, Saebyul Park, Sangeon Yong, and Juhan Nam, · 2020
Cited alongside, same era.
Closest in time.