Fetching the paper…
Reading the bibliography…
W. Hartmann,
2004
Earlier work this paper cites.
K. Saino, H. Zen, Y. Nankaku, A. Lee, and K. Tokuda, “An hmm-based singing voice synthesis system,” in
2006
Earlier work this paper cites.
K. Oura, A. Mase, T. Yamada, S. Muto, Y. Nankaku, and K. Tokuda, “Recent development of the hmm-based singing voice synthesis system—sinsy,” in
2010
Earlier work this paper cites.
F. Villavicencio and J. Bonada, “Applying voice conversion to concatenative singing-voice synthesis.” in
2010
Earlier work this paper cites.
J. C. Smith, “Correlation analyses of encoded music performance,” Ph.D. dissertation, Stanford, CA, USA, 2013
2013
Earlier work this paper cites.
Z. Duan, H. Fang, B. Li, K. C. Sim, and Y. Wang, “The nus sung and spoken lyrics corpus: A quantitative comparison of singing and speech,” in
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” in
2014
Earlier work this paper cites.
——, “Statistical singing voice conversion with direct waveform modification based on the spectrum differential,” in
2014
Earlier work this paper cites.
K. Kobayashi, T. Toda, G. Neubig, S. Sakti, and S. Nakamura, “Statistical singing voice conversion based on direct waveform modification with global variance,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
J. Bonada, M. Umbert, and M. Blaauw, “Expressive singing synthesis based on unit selection for the singing synthesis challenge 2016,” in
2016
Cited alongside, same era.
Y. Ganin
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Cited alongside, same era.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with WaveNet autoencoders,” in
2017
Cited alongside, same era.
D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (elus),” in
2017
Later among the works it cites.
C.-i. Wang and G. Tzanetakis, “Singing style investigation by residual siamese convolutional neural networks,” in
2018
Later among the works it cites.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop,” in
2018
Later among the works it cites.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,”
2018
Later among the works it cites.
N. Mor, L. Wolf, A. Polyak, and Y. Taigman, “A universal music translation network,” in
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural Discrete Representation Learning,” in
2017
Cited alongside, same era.
——, “Voice Conversion from Unaligned Corpora Using Variational Autoencoding Wasserstein Generative Adversarial Networks,” in
2017
Cited alongside, same era.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein Generative Adversarial Networks,” in
2017
Cited alongside, same era.
M. Blaauw and J. Bonada, “A neural parametric singing synthesizer modeling timbre and expression from natural songs,”
2017
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,”
2017
Cited alongside, same era.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie, C. Gulcehre, A. van den Oord, O. Vinyals, and N. de Freitas, “Sample efficient adaptive text-to-speech,” in
2019
Closest in time.
A. Polyak and L. Wolf, “Attention-based wavenet autoencoder for universal voice conversion,” in
2019
Closest in time.
M. Blaauw, J. Bonada, and R. Daido, “Data efficient voice cloning for neural singing synthesis,” in
2019
Closest in time.
X. Chen, W. Chu, J. Guo, and N. Xu, “Singing voice conversion with non-parallel data,”
2019
Closest in time.
E. Nachmani and L. Wolf, “Unsupervised polyglot text to speech,” in
2019
Closest in time.