Fetching the paper…
Reading the bibliography…
Fast and user-controllable music generation could enable novel ways of composing or performing music.
J. Liu, Y. Chen, Y. Yeh, and Y. Yang, “Unconditional audio generation with generative adversarial networks and cycle regularization,” in 21st Annual Conference of the International Speech Communication Association (INTERSPEECH) , Oct. 2020, pp. 1997–2001
2001
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in 2nd International Conference on Learning Representations (ICLR) , Apr. 2014
2014
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27 , Dec. 2014, pp. 2672–2680
2014
Earlier work this paper cites.
J. Schlüter and S. Böck, “Improved musical onset detection with convolutional neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2014, pp. 6979–6983
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations (ICLR) , May 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning (ICML) , ser. JMLR Workshop and Conference Proceedings, vol. 37, Jul. 2015, pp. 448–456
2015
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” in The 9th ISCA Speech Synthesis Workshop , Sep. 2016, p. 125
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Böck, F. Korzeniowski, J. Schlüter, F. Krebs, and G. Widmer, “madmom: A new python audio and music signal processing library,” in Proceedings of the 2016 ACM Conference on Multimedia Conference (MM) , Oct. 2016, pp. 1174–1178
2016
Earlier work this paper cites.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. C. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” in 5th International Conference on Learning Representations (ICLR) , Apr. 2017
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Advances in Neural Information Processing Systems 30 , Dec. 2017, pp. 6306–6315
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems 30 , Dec. 2017, pp. 5998–6008
2017
Earlier work this paper cites.
J. H. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with WaveNet autoencoders,” in Proceedings of the 34th International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 70, Aug. 2017, pp. 1068–1077
2017
Earlier work this paper cites.
J. H. Lim and J. C. Ye, “Geometric GAN,” arXiv preprint arXiv:1705.02894 , 2017
2017
Earlier work this paper cites.
X. Huang and S. J. Belongie, “Arbitrary style transfer in real-time with adaptive instance normalization,” in IEEE International Conference on Computer Vision (ICCV) , Oct. 2017, pp. 1510–1519
2017
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of GANs for improved quality, stability, and variation,” in 6th International Conference on Learning Representations (ICLR) , Apr. 2018
2018
Cited alongside, same era.
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, “Spectral normalization for generative adversarial networks,” in 6th International Conference on Learning Representations (ICLR) , Apr. 2018
2018
Cited alongside, same era.
L. M. Mescheder, A. Geiger, and S. Nowozin, “Which training methods for GANs do actually converge?” in Proceedings of the 35th International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 80, Jul. 2018, pp. 3478–3487
2018
Cited alongside, same era.
H. Schreiber and M. Müller, “A single-step approach to musical tempo estimation using a convolutional neural network,” in Proceedings of the 19th International Society for Music Information Retrieval Conference (ISMIR) , Sep. 2018, pp. 98–105
2018
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems 33 , Dec. 2020
2020
Later among the works it cites.
J. Nistal, S. Lattner, and G. Richard, “DRUMGAN: synthesis of drum sounds with timbral feature conditioning using generative adversarial networks,” in Proceedings of the 21th International Society for Music Information Retrieval Conference (ISMIR) , Oct. 2020, pp. 590–597
2020
Later among the works it cites.
J. Nistal, S. Lattner, and G. Richard, “Comparing representations for audio synthesis using generative adversarial networks,” in 28th European Signal Processing Conference (EUSIPCO) . IEEE, Jan. 2020, pp. 161–165
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A. Razavi, A. van den Oord, and O. Vinyals, “Generating diverse high-fidelity images with VQ-VAE-2,” in Advances in Neural Information Processing Systems 32 , Dec. 2019, pp. 14 837–14 847
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Advances in Neural Information Processing Systems 32 , Dec. 2019, pp. 14 881–14 892
2019
Cited alongside, same era.
C. Donahue, J. J. McAuley, and M. S. Puckette, “Adversarial audio synthesis,” in 7th International Conference on Learning Representations (ICLR) , May 2019
2019
Cited alongside, same era.
J. H. Engel, K. K. Agrawal, S. Chen, I. Gulrajani, C. Donahue, and A. Roberts, “GANSynth: Adversarial neural audio synthesis,” in 7th International Conference on Learning Representations (ICLR) , May 2019
2019
Cited alongside, same era.
O. Prykhodko, S. Johansson, P. Kotsias, J. Arús-Pous, E. J. Bjerrum, O. Engkvist, and H. Chen, “A de novo molecular generation method using latent vector based generative adversarial network,” J. Cheminformatics , vol. 11, no. 1, p. 74, 2019
2019
Cited alongside, same era.
C. H. Lin, C. Chang, Y. Chen, D. Juan, W. Wei, and H. Chen, “COCO-GAN: generation by parts via conditional coordinating,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Oct. 2019, pp. 4511–4520
2019
Cited alongside, same era.
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in 20th Annual Conference of the International Speech Communication Association (INTERSPEECH) , Sep. 2019, pp. 1526–1530
2019
Cited alongside, same era.
J. H. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: differentiable digital signal processing,” in 8th International Conference on Learning Representations (ICLR) , Apr. 2020
2020
Later among the works it cites.
2021
Later among the works it cites.
I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R. J. Skerry-Ryan, and Y. Wu, “Parallel tacotron 2: A non-autoregressive neural TTS model with differentiable duration modeling,” in 22nd Annual Conference of the International Speech Communication Association (INTERSPEECH) , Aug. 2021, pp. 141–145
2021
Later among the works it cites.
2021
Later among the works it cites.
I. Skorokhodov, G. Sotnikov, and M. Elhoseiny, “Aligning latent and image spaces to connect the unconnectable,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Oct. 2021, pp. 14 124–14 133
2021
Later among the works it cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, Jun. 2021, pp. 12 873–12 883
2021
Later among the works it cites.
B. Liu, Y. Zhu, K. Song, and A. Elgammal, “Towards faster and stabilized GAN training for high-fidelity few-shot image synthesis,” in 9th International Conference on Learning Representations (ICLR) , May 2021
2021
Later among the works it cites.
A. Sauer, K. Chitta, J. Müller, and A. Geiger, “Projected GANs converge faster,” in Advances in Neural Information Processing Systems 34 , Dec. 2021, pp. 17 480–17 492
2021
Later among the works it cites.
C. H. Lin, Y.-C. Cheng, H.-Y. Lee, S. Tulyakov, and M.-H. Yang, “InfinityGAN: Towards infinite-pixel image synthesis,” in 10th International Conference on Learning Representations (ICLR) , Apr. 2022
2022
Closest in time.
T. Kaneko, K. Tanaka, H. Kameoka, and S. Seki, “iSTFTNET: fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, May 2022, pp. 6207–6211
2022
Closest in time.