Fetching the paper…
Reading the bibliography…
In a recent paper, we have presented a generative adversarial network (GAN)-based model for unconditional generation of the mel-spectrograms of singing voices.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
2007
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
M. Mirza and S. Osindero, “Conditional generative adversarial nets,”
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder-decoder approaches,” in
2014
Earlier work this paper cites.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” in
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
A. v. d. Oord, N. Kalchbrenner, and K. Kavukcuoglu, “Pixel recurrent neural networks,” in
2016
Earlier work this paper cites.
A. v. d. Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves, and K. Kavukcuoglu, “Conditional image generation with PixelCNN decoders,” in
2016
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, and X. Chen, “Improved techniques for training GANs,” in
2016
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,” in
2017
Earlier work this paper cites.
D. Berthelot, T. Schumm, and L. Metz, “BEGAN: Boundary equilibrium generative adversarial networks,”
2017
Earlier work this paper cites.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,”
2017
Earlier work this paper cites.
K. Ito, “The LJ speech dataset,”
2017
Earlier work this paper cites.
H.-W. Dong, W.-Y. Hsiao, L.-C. Yang, and Y.-H. Yang, “MuseGAN: Symbolic-domain music generation and accompaniment with multi-track sequential generative adversarial networks,” in
2018
Cited alongside, same era.
S. Dieleman, A. van den Oord, and K. Simonyan, “The challenge of realistic music generation: modelling raw audio at scale,”
2018
Cited alongside, same era.
D. Ulyanov, A. Vedaldi, and V. Lempitsky, “It takes (only) two: Adversarial generator-encoder networks,” Tech. Rep., apr 2018
2018
Cited alongside, same era.
E. Richardson and Y. Weiss, “On GANs and GMMs,” in
2018
Cited alongside, same era.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of StyleGAN,”
2019
Cited alongside, same era.
2019
Later among the works it cites.
S. Vasquez and M. Lewis, “MelNet: A generative model for audio in the frequency domain,”
2019
Later among the works it cites.
Q. Mao, H.-Y. Lee, H.-Y. Tseng, S. Ma, and M.-H. Yang, “Mode seeking generative adversarial networks for diverse image synthesis,” in
2019
Later among the works it cites.
J.-Y. Liu and Y.-H. Yang, “Dilated convolution with dilated GRU for music source separation,” in
2019
Later among the works it cites.
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music Transformer: Generating music with long-term structure,” in
2019
Cited alongside, same era.
J. Engel, K. K. Agrawal, S. Chen, I. Gulrajani, C. Donahue, and A. Roberts, “GANsynth: Adversarial neural audio synthesis,” in
2019
Cited alongside, same era.
B. Wang and Y.-H. Yang, “PerformanceNet: Score-to-audio music generation with multi-band convolutional residual network,” in
2019
Cited alongside, same era.
E. Dunbar, R. Algayres, J. Karadayi, M. Bernard, J. Benjumea, X. Cao, L. Miskic, C. Dugrain, L. Ondel, A. W. Black, L. Besacier, S. Sakti, and E. Dupoux, “The zero resource speech challenge 2019: TTS without T,” in
2019
Cited alongside, same era.
K.-Y. Chen, C.-P. Tsai, D.-R. Liu, H.-Y. Lee, and L.-S. Lee, “Completely unsupervised phoneme recognition by a generative adversarial network harmonized with iteratively refined hidden Markov models,” in
2019
Cited alongside, same era.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,”
2019
Cited alongside, same era.
R. Eloff, A. Nortje, B. van Niekerk, A. Govender, L. Nortje, A. Pretorius, E. V. Biljon, E. van der Westhuizen, L. van Staden, and H. Kamper, “Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks,” in
2019
Cited alongside, same era.
2019
Later among the works it cites.
S. Kum and J. Nam, “Joint detection and classification of singing voice melody using convolutional recurrent neural networks,”
2019
Later among the works it cites.
J. Lee, H. Choi, C. Jeon, J. Koo, and K. Lee, “Adversarially trained end-to-end Korean singing voice synthesis system,” in
2019
Later among the works it cites.
Y.-S. Huang and Y.-H. Yang, “Pop music transformer: Generating music with rhythm and harmony,”
2020
Closest in time.
J. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: Differentiable digital signal processing,” in
2020
Closest in time.
J. Parekh, P. Rao, and Y.-H. Yang, “Speech-to-singing conversion in an encoder-decoder framework,” in
2020
Closest in time.
J.-Y. Liu, Y.-H. Chen, Y.-C. Yeh, and Y.-H. Yang, “Score and lyrics-free singing voice generation,” in
2020
Closest in time.