Fetching the paper…
Reading the bibliography…
Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score.
Singing voice synthesis based on convolutional neural networks
Nakamura, K.; Hashimoto, K.; Oura, K.; Nankaku, Y.; and Tokuda, K. 2019 · 1904
Earlier work this paper cites.
Deepsinger: Singing voice synthesis with data mined from the web
Ren, Y.; Tan, X.; Qin, T.; Luan, J.; Zhao, Z.; and Liu, T.-Y. 2020 · 1989
Earlier work this paper cites.
Using Cyclic Noise as the Source Signal for Neural Source-Filter-Based Speech Waveform Model
Wang, X.; and Yamagishi, J. 2020 · 1996
Earlier work this paper cites.
Concatenation-based MIDI-to-singing voice synthesis
Macon, M.; Jensen-Link, L.; George, E. B.; Oliverio, J.; and Clements, M. 1997 · 1997
Earlier work this paper cites.
Gu, Y.; Yin, X.; Rao, Y.; Wan, Y.; Tang, B.; Zhang, Y.; Chen, J.; Wang, Y.; and Ma, Z. 2020 · 2004
Earlier work this paper cites.
An HMM-based singing voice synthesis system
Saino, K.; Zen, H.; Nankaku, Y.; Lee, A.; and Tokuda, K. 2006 · 2006
Earlier work this paper cites.
Vocaloid-commercial singing synthesizer based on sample concatenation
Kenmochi, H.; and Ohshita, H. 2007 · 2007
Earlier work this paper cites.
HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
Chen, J.; Tan, X.; Luan, J.; Qin, T.; and Liu, T.-Y. 2020 · 2009
Earlier work this paper cites.
Recent development of the HMM-based singing voice synthesis system—Sinsy
Oura, K.; Mase, A.; Yamada, T.; Muto, S.; Nankaku, Y.; and Tokuda, K. 2010 · 2010
Earlier work this paper cites.
A Connection Between Score Matching and Denoising Autoencoders
Vincent, P. 2011 · 2011
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Morise, M.; Yokomori, F.; and Ozawa, K. 2016 · 2016
Earlier work this paper cites.
Singing Voice Synthesis Based on Deep Neural Networks
Nishimura, M.; Hashimoto, K.; Oura, K.; Nankaku, Y.; and Tokuda, K. 2016 · 2016
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
Oord, A. v. d.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
A neural parametric singing synthesizer modeling timbre and expression from natural songs
Blaauw, M.; and Bonada, J. 2017 · 2017
Cited alongside, same era.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi
McAuliffe, M.; Socolof, M.; Mihuc, S.; Wagner, M.; and Sonderegger, M. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Korean Singing Voice Synthesis System based on an LSTM Recurrent Neural Network
Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
Kim, J.; Kim, S.; Kong, J.; and Yoon, S. 2020 · 2020
Later among the works it cites.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Kong, J.; Kim, J.; and Bae, J. 2020 · 2020
Later among the works it cites.
Adversarially Trained Multi-Singer Sequence-to-Sequence Singing Synthesizer
Wu, J.; and Luan, J. 2020 · 2020
Later among the works it cites.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Yamamoto, R.; Song, E.; and Kim, J.-M. 2020 · 2020
Later among the works it cites.
DurIAN-SC: Duration Informed Attention Network Based Singing Voice Conversion System
Zhang, L.; Yu, C.; Lu, H.; Weng, C.; Zhang, C.; Wu, Y.; Xie, X.; Li, Z.; and Yu, D. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kim, J.; Choi, H.; Park, J.; Kim, S.; Kim, J.; and Hahn, M. 2018 · 2018
Cited alongside, same era.
A wavenet for speech denoising
Rethage, D.; Pons, J.; and Serra, X. 2018 · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Shen, J.; Pang, R.; Weiss, R. J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; et al. 2018 · 2018
Cited alongside, same era.
The LJ Speech Dataset
Ito, K.; and Johnson, L. 2017 · 2019
Cited alongside, same era.
Adversarially Trained End-to-End Korean Singing Voice Synthesis System
Lee, J.; Choi, H.-S.; Jeon, C.-B.; Koo, J.; and Lee, K. 2019 · 2019
Cited alongside, same era.
Generative Modeling by Estimating Gradients of the Data Distribution
Song, Y.; and Ermon, S. 2019 · 2019
Cited alongside, same era.
Sequence-to-sequence singing synthesis using the feed-forward transformer
Blaauw, M.; and Bonada, J. 2020 · 2020
Cited alongside, same era.
WaveGrad: Estimating Gradients for Waveform Generation
Chen, N.; Zhang, Y.; Zen, H.; Weiss, R. J.; Norouzi, M.; and Chan, W. 2021 · 2021
Closest in time.
Diff-tts: A denoising diffusion model for text-to-speech
Jeong, M.; Kim, H.; Cheon, S. J.; Choi, B. J.; and Kim, N. S. 2021 · 2021
Closest in time.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; and Catanzaro, B. 2021 · 2021
Closest in time.
Improved denoising diffusion probabilistic models
Nichol, A. Q.; and Dhariwal, P. 2021 · 2021
Closest in time.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T.-Y. 2021 · 2021
Closest in time.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2021 · 2021
Closest in time.
Score-Based Generative Modeling through Stochastic Differential Equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2021 · 2021
Closest in time.