Fetching the paper…
Reading the bibliography…
We present UTACO, a singing synthesis model based on an attention-based sequence-to-sequence mechanism and a vocoder based on dilated causal convolutions.
P. R. Cook, “Singing voice synthesis: History, current work, and future directions,”
1996
Earlier work this paper cites.
M. Good, “MusicXML in Commercial Applications | Songs and Schemas,”
2006
Earlier work this paper cites.
H. Kenmochi and Hayato Ohshita, “Vocaloid-commercial singing synthesizer based on sample concatenation.” in
2007
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,”
2014
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio Augmentation for Speech Recognition,” in
2015
Earlier work this paper cites.
I. R. Assembly, “ITU-R BS. 1534-3: Method for the subjective assessment of intermediate quality level of audio systems. October 2015.”
2015
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” in
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Earlier work this paper cites.
M. Blaauw and J. Bonada, “A neural parametric singing synthesizer,” 2017, pp. 4001–4005
2017
Earlier work this paper cites.
——, “A neural parametric singing synthesizer modeling timbre and expression from natural songs,”
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Hono, S. Murata, K. Nakamura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Recent Development of the DNN-based Singing Voice Synthesis System-Sinsy,” in
2018
Cited alongside, same era.
T. Merritt, B. Putrycz, A. Nadolski, T. Ye, D. Korzekwa, W. Dolecki, T. Drugman, V. Klimkov, A. Moinet, A. Breen
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Closest in time.
F. Bous and A. Roebel, “Analysing deep learning-spectral envelope prediction methods for singing synthesis,” in
2019
Closest in time.
P. Chandna, M. Blaauw, J. Bonada, and E. Gómez, “WGansing: A multi-voice singing voice synthesizer based on the Wasserstein-Gan,” in
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Y. Wada, R. Nishikimi, E. Nakamura, K. Itoyama, and K. Yoshii, “Sequential Generation of Singing F0 Contours from Musical Note Sequences Based on WaveNet,” in
2018
Cited alongside, same era.
K. Hua, “Modeling Singing F0 With Neural Network Driven Transition-Sustain Models,”
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,” in
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in
2018
Cited alongside, same era.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in
2018
Cited alongside, same era.
Closest in time.
V. Klimkov, S. Ronanki, J. Rohnke, and T. Drugman, “Fine-grained robust prosody transfer for single-speaker neural text-to-speech,”
2019
Closest in time.
2019
Closest in time.
J. Latorre, J. Lachowicz, J. Lorenzo-Trueba, T. Merritt, T. Drugman, S. Ronanki, and V. Klimkov, “Effect of data reduction on sequence-to-sequence neural tts,” in
2019
Closest in time.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in
2019
Closest in time.