Fetching the paper…
Reading the bibliography…
We introduce Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts.
Evaluation of algorithms using games: The case of music tagging
Law, E., West, K., Mandel, M. I., Bay, M., and Downie, J. S · 2009
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark, 2016
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R., and Wilson, K · 2017
Earlier work this paper cites.
Fr \ \backslash ’echet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2018
Earlier work this paper cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Wang, Y., Stanton, D., Zhang, Y., Ryan, R.-S., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R. A · 2018
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al · 2020
Earlier work this paper cites.
Towards Learning a Universal Non-Semantic Representation of Speech
Shor, J., Jansen, A., Maor, R., Lang, O., Tuval, O., de Chaumont Quitry, F., Tagliasacchi, M., Shavitt, I., Emanuel, D., and Haviv, Y · 2020
Earlier work this paper cites.
From artificial neural networks to deep learning for music generation: history, concepts and trends
Briot, J.-P · 2021
Cited alongside, same era.
Wavegrad: Estimating gradients for waveform generation
Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W · 2021
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Cited alongside, same era.
Grad-tts: A diffusion probabilistic model for text-to-speech
Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Later among the works it cites.
Mulan: A joint embedding of music audio and natural language
Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. W · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Later among the works it cites.
Contrastive audio-language learning for music, 2022
Manco, I., Benetos, E., Quinton, E., and Fazekas, G · 2022
Later among the works it cites.
https://github.com/mubertai/mubert-text-to-music
MubertAI · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Cited alongside, same era.
Wu, S. and Shi, Z · 2021
Cited alongside, same era.
GSPMD: general and scalable parallelization for ML computation graphs
Xu, Y., Lee, H., Chen, D., Hechtman, B. A., Huang, Y., Joshi, R., Krikun, M., Lepikhin, D., Ly, A., Maggioni, M., Pang, R., Shazeer, N., Wang, S., Wang, T., Wu, Y., and Chen, Z · 2021
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation, 2022
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Cited alongside, same era.
Infergrad: Improving diffusion models for vocoder by considering inference in training
Chen, Z., Tan, X., Wang, K., Pan, S., Mandic, D., He, L., and Zhao, S · 2022
Cited alongside, same era.
Riffusion - Stable diffusion for real-time music generation
Forsgren, S. and Martiros, H · 2022
Cited alongside, same era.
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
Palette: Image-to-image diffusion models
Saharia, C., Chan, W., Chang, H., Lee, C., Ho, J., Salimans, T., Fleet, D., and Norouzi, M · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Later among the works it cites.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., and Frank, C · 2023
Closest in time.
Moûsai: Text-to-music generation with long-context latent diffusion, 2023
Schneider, F., Jin, Z., and Schölkopf, B · 2023
Closest in time.