Fetching the paper…
Reading the bibliography…
Recent advancements in deep generative models present new opportunities for music production but also pose challenges, such as high computational demands and limited audio quality.
I. J. Goodfellow, J. Pouget-Abadie et al. , “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27 , Dec. 2014, pp. 2672–2680
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in 2nd International Conference on Learning Representations (ICLR) , Apr. 2014
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. A. Weiss et al. , “Deep unsupervised learning using nonequilibrium thermodynamics,” in Proceedings of the 32nd International Conference on Machine Learning, ICML , 2015
2015
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” in Proc. of the 9th ISCA Speech Synthesis Workshop , 2016
2016
Earlier work this paper cites.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. C. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” in Proc. of 5th International Conference on Learning Representations, ICLR , 2017
2017
Earlier work this paper cites.
M. Binkowski, D. J. Sutherland, M. Arbel, and A. Gretton, “Demystifying MMD gans,” in Proc. of the 6th International Conference on Learning Representations, ICLR , 2018
2018
Earlier work this paper cites.
S. Lattner and M. Grachten, “High-level control of drum track generation using learned patterns of rhythmic interaction,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA . IEEE, 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in 7th International Conference on Learning Representations, ICLR , 2019
2019
Earlier work this paper cites.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms,” in Proc. of the 20th Conference of the International Speech Communication Association (InterSpeech) , 2019
2019
Earlier work this paper cites.
J. Nistal, S. Lattner, and G. Richard, “DrumGAN: Synthesis of drum sounds with timbral feature conditioning,” in Proc. of the 21st International Society for Music Information Retrieval Conference, ISMIR , 2020
2020
Earlier work this paper cites.
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, “Jukebox: A generative model for music,” in CoRR , 2020
2020
Earlier work this paper cites.
M. Grachten, S. Lattner, and E. Deruty, “BassNet: A variational gated autoencoder for conditional generation of bass guitar tracks with learned interactive control,” in Applied Sciences , 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems, NeurIPS , 2020
2020
Earlier work this paper cites.
M. F. Naeem, S. J. Oh, Y. Uh, Y. Choi, and J. Yoo, “Reliable fidelity and diversity metrics for generative models,” in Proc. of the 37th International Conference on Machine Learning, ICML , 2020
2020
Earlier work this paper cites.
A. Caillon and P. Esling, “RAVE: A variational autoencoder for fast and high-quality neural audio synthesis,” in CoRR , 2021
2021
Earlier work this paper cites.
S. Rouard and G. Hadjeres, “CRASH: raw audio score-based generative modeling for controllable high-resolution drum sound synthesis,” in Proc. of the 22nd International Society for Music Information Retrieval Conference, ISMIR , 2021
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in Proc. of the 9th International Conference on Learning Representations, ICLR , 2021
2021
Earlier work this paper cites.
T. Karras, M. Aittala, T. Aila, and S. Laine, “Score-based generative modeling through stochastic differential equations,” in Proc. of the International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
D. Barry, Q. Zhang, P. W. Sun, and A. Hines, “Go listen: An end-to-end online listening test platform.” Journal of Open Research Software , vol. 9, no. 1, 2021
2021
Cited alongside, same era.
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” in Proc. of the 36th Conference on Neural Information Processing Systems (NeurIPS) , 2022
2022
Cited alongside, same era.
M. Pasini and J. Schlüter, “Musika! Fast infinite waveform music generation,” in Proc. of the 23rd International Society for Music Information Retrieval Conference, ISMIR , 2022
2022
Cited alongside, same era.
J. Nistal, C. Aouameur, I. Velarde, and S. Lattner, “DrumGAN VST: A plugin for drum sound analysis/synthesis with autoencoding generative adversarial networks,” in Proc. of International Conference on Machine Learning ICML, Workshop on Machine Learning for Audio Synthesis, MLAS , 2022
2022
Cited alongside, same era.
H. F. García, P. Seetharaman, R. Kumar, and B. Pardo, “VampNet: Music generation via masked acoustic token modeling,” in Proc. of the 24th International Society for Music Information Retrieval Conference, ISMIR , 2023
2023
Later among the works it cites.
F. Schneider, O. Kamal, Z. Jin, and B. Schölkopf, “Moûsai: Text-to-music generation with long-context latent diffusion,” in CoRR , 2023
2023
Later among the works it cites.
P. Li, B. Chen, Y. Yao, Y. Wang, A. Wang, and A. Wang, “JEN-1: text-guided universal music generation with omnidirectional,” in CoRR , 2023
2023
Later among the works it cites.
S.-L. Wu, C. Donahue, S. Watanabe, and N. J. Bryan, “Music controlnet: Multiple time-varying controls for music generation,” in CoRR , 2023
2023
Later among the works it cites.
M. Levy, B. D. Giorgi, F. Weers, A. Katharopoulos, and T. Nickson, “Controllable music production with diffusion models and guidance gradients,” CoRR , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” in IEEE ACM Trans. Audio Speech Lang. Process. , 2022
2022
Cited alongside, same era.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” in CoRR , 2022
2022
Cited alongside, same era.
Y. Wu, C. Chiu, and Y. Yang, “Jukedrummer: Conditional beat-aware audio-domain drum accompaniment generation via transformer VQ-VAE,” in Proc. of the 3rd International Society for Music Information Retrieval Conference, ISMIR , 2022
2022
Cited alongside, same era.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022
Cited alongside, same era.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman, “Make-A-Video: Text-to-video generation without text-video data,” in Proc. of the 11th International Conference on Learning Representations, ICLR , 2023
2023
Cited alongside, same era.
J. Betker, G. Goh, L. Jing, TimBrooks, J. Wang, L. Li, LongOuyang, JuntangZhuang, JoyceLee, YufeiGuo, WesamManassra, PrafullaDhariwal, CaseyChu, YunxinJiao, and A. Ramesh, “Improving image generation with better captions,” in CoRR , 2023
2023
Cited alongside, same era.
A. Agostinelli, T. I. Denk, Z. Borsos, J. H. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. H. Frank, “MusicLM: Generating Music From Text,” in CoRR , 2023
2023
Cited alongside, same era.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” in Advances in Neural Information Processing Systems 36 NeurIPS , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
C. Donahue, A. Caillon, A. Roberts, E. Manilow, P. Esling, A. Agostinelli, M. Verzetti, I. Simon, O. Pietquin, N. Zeghidour, and J. H. Engel, “Singsong: Generating musical accompaniments from singing,” in CoRR , 2023
2023
Later among the works it cites.
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, “Consistency models,” in International Conference on Machine Learning, ICML , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,” in Proc. of the International Conference on Machine Learning, ICML , 2023
2023
Later among the works it cites.
Z. Evans, C. Carr, J. Taylor, S. H. Hawley, and J. Pons, “Fast timing-conditioned latent audio diffusion,” in CoRR , 2024
2024
Closest in time.
M. Pasini, S. Lattner, and G. Fazekas, “Music2latent: Consistency autoencoders for latent audio compression,” in Proc. of the International Society for Music Information Retrieval (ISMIR) , 2024
2024
Closest in time.
M. Pasini, M. Grachten, and S. Lattner, “Bass accompaniment generation via latent diffusion,” in IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP . IEEE, 2024
2024
Closest in time.
A. Ziv, I. Gat, G. L. Lan, T. Remez, F. Kreuk, A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “Masked audio generation using a single non-autoregressive transformer,” in CoRR , 2024
2024
Closest in time.
Y. Zhang, Y. Ikemiya, G. Xia, N. Murata, M. A. M. Ramírez, W. Liao, Y. Mitsufuji, and S. Dixon, “Musicmagus: Zero-shot text-to-music editing via diffusion models,” CoRR , 2024
2024
Closest in time.
H. Manor and T. Michaeli, “Zero-shot unsupervised and text-based audio editing using DDPM inversion,” CoRR , 2024
2024
Closest in time.
Z. Novack, J. J. McAuley, T. Berg-Kirkpatrick, and N. J. Bryan, “DITTO: diffusion inference-time t-optimization for music generation,” CoRR , 2024
2024
Closest in time.
E. Postolache, G. Mariani, L. Cosmo, E. Benetos, and E. Rodolà, “Generalized multi-source inference for text conditioned music diffusion models,” CoRR , 2024
2024
Closest in time.
M. Grachten and J. Nistal, “Audio Prompt Adherence: A measure for evaluating musical accompaniment systems,” in CoRR , 2024
2024
Closest in time.