Fetching the paper…
Reading the bibliography…
Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts.
J. Serra, M. Müller, P. Grosche, and J. L. Arcos, “Unsupervised music structure annotation by time series structure features and segment similarity,” IEEE Trans. on Multimedia , 2014
2014
Earlier work this paper cites.
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” arXiv , 2016
2016
Earlier work this paper cites.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” arXiv , 2016
2016
Earlier work this paper cites.
T. Chen, B. Xu, C. Zhang, and C. Guestrin, “Training deep nets with sublinear memory cost,” arXiv , 2016
2016
Earlier work this paper cites.
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, “Fma: A dataset for music analysis,” arXiv , 2016
2016
Earlier work this paper cites.
M. Schoeffler, S. Bartoschek, F.-R. Stöter, M. Roess, S. Westphal, B. Edler, and J. Herre, “webMUSHRA—a comprehensive framework for web-based listening tests,” Journal of Open Research Software , 2018
2018
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv , 2019
2019
Earlier work this paper cites.
A. L. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, “Look, listen, and learn more: Design choices for deep audio embeddings,” IEEE Int. Conf. on Acoustics, Speech and Sig. Proc. (ICASSP) , 2019
2019
Earlier work this paper cites.
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, “Jukebox: A generative model for music,” arXiv , 2020
2020
Earlier work this paper cites.
L. Ziyin, T. Hartwig, and M. Ueda, “Neural networks fail to learn periodic functions and how to fix it,” arXiv , 2020
2020
Earlier work this paper cites.
C. J. Steinmetz and J. D. Reiss, “auraloss: Audio focused loss functions in PyTorch,” in Digital Music Research Network One-day Workshop (DMRN+15) , 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” arXiv , 2020
2020
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Trans. on Audio, Speech, and Language Processing , 2021
2021
Earlier work this paper cites.
S. Rouard and G. Hadjeres, “CRASH: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis,” arXiv , 2021
2021
Earlier work this paper cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” arXiv , 2022
2022
Earlier work this paper cites.
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, “Clap: Learning audio concepts from natural language supervision,” arXiv , 2022
2022
Earlier work this paper cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and memory-efficient exact attention with IO-awareness,” arXiv , 2022
2022
Cited alongside, same era.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in IEEE Int. Conf. on Acoustics, Speech and Sig. Proc. (ICASSP) , 2022
2022
Cited alongside, same era.
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,” arXiv , 2022
2022
Cited alongside, same era.
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “DPM-solver++: Fast solver for guided sampling of diffusion probabilistic models,” arXiv , 2022
2022
Cited alongside, same era.
K. Koutini, J. Schlüter, H. Eghbal-zadeh, and G. Widmer, “Efficient training of audio transformers with patchout,” Conf. of the Int. Speech Comm. Assoc. (INTERSPEECH) , 2022
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” IEEE Int. Conf. on Acoustics, Speech and Sig. Proc. (ICASSP) , 2023
2023
Later among the works it cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proc. of the IEEE/CVF Int. Conf. on Comp. Vision (ICCV) , 2023
2023
Later among the works it cites.
M. Levy, B. Di Giorgi, F. Weers, A. Katharopoulos, and T. Nickson, “Controllable music production with diffusion models and guidance gradients,” arXiv , 2023
2023
Later among the works it cites.
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” arXiv , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank, “MusicLM: Generating music from text,” arXiv , 2023
2023
Cited alongside, same era.
F. Schneider, Z. Jin, and B. Schölkopf, “Moûsai: Text-to-music generation with long-context latent diffusion,” arXiv , 2023
2023
Cited alongside, same era.
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank, J. Engel, Q. V. Le, W. Chan, Z. Chen, and W. Han, “Noise2music: Text-conditioned music generation with diffusion models,” arXiv , 2023
2023
Cited alongside, same era.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” arXiv , 2023
2023
Cited alongside, same era.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” arXiv , 2023
2023
Cited alongside, same era.
H. F. Garcia, P. Seetharaman, R. Kumar, and B. Pardo, “VampNet: Music generation via masked acoustic token modeling,” arXiv , 2023
2023
Cited alongside, same era.
M. W. Y. Lam, Q. Tian, T. Li, Z. Yin, S. Feng, M. Tu, Y. Ji, R. Xia, M. Ma, X. Song, J. Chen, Y. Wang, and Y. Wang, “Efficient neural music generation,” arXiv , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
I. Manco, B. Weck, S. Doh, M. Won, Y. Zhang, D. Bogdanov, Y. Wu, K. Chen, P. Tovstogan, E. Benetos, E. Quinton, G. Fazekas, and J. Nam, “The Song Describer Dataset: a corpus of audio captions for music-and-language evaluation,” arXiv , 2023
2023
Later among the works it cites.
N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramer, B. Balle, D. Ippolito, and E. Wallace, “Extracting training data from diffusion models,” in USENIX Security Symposium , 2023
2023
Later among the works it cites.
A. Ziv, I. Gat, G. L. Lan, T. Remez, F. Kreuk, A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “Masked audio generation using a single non-autoregressive transformer,” arXiv , 2024
2024
Closest in time.
G. Cideron, S. Girgin, M. Verzetti, D. Vincent, M. Kastelic, Z. Borsos, B. McWilliams, V. Ungureanu, O. Bachem, O. Pietquin, M. Geist, L. Hussenot, N. Zeghidour, and A. Agostinelli, “MusicRL: Aligning music generation to human preferences,” arXiv , 2024
2024
Closest in time.
Z. Evans, C. Carr, J. Taylor, S. Hawley, and J. Pons, “Fast timing-conditioned latent audio diffusion,” arXiv , 2024
2024
Closest in time.
J. Parker, J. Spijkervet, K. Kosta, F. Yesiler, B. Kuznetsov, J.-C. Wang, M. Avent, J. Chen, and D. Le, “StemGen: A music generation model that listens,” IEEE Int. Conf. on Acoustics, Speech and Sig. Proc. (ICASSP) , 2024
2024
Closest in time.
Z. Novack, J. McAuley, T. Berg-Kirkpatrick, and N. J. Bryan, “Ditto: Diffusion inference-time t-optimization for music generation,” arXiv , 2024
2024
Closest in time.
M. Pasini, M. Grachten, and S. Lattner, “Bass accompaniment generation via latent diffusion,” arXiv , 2024
2024
Closest in time.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems (NeurIPS) , 2024
2024
Closest in time.
S. Lin, B. Liu, J. Li, and X. Yang, “Common diffusion noise schedules and sample steps are flawed,” IEEE/CVF Winter Conf. on Applications of Comp. Vision , 2024
2024
Closest in time.
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel et al. , “Scaling rectified flow transformers for high-resolution image synthesis,” arXiv , 2024
2024
Closest in time.