Fetching the paper…
Reading the bibliography…
Generating long-form 44.1kHz stereo audio from text prompts can be computationally demanding.
Loudness, Its Definition, Measurement and Calculation
Fletcher, H. and Munson, W. A · 2005
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., et al · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A · 2017
Earlier work this paper cites.
Adversarial audio synthesis
Donahue, C., McAuley, J., and Puckette, M · 2018
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2018
Earlier work this paper cites.
Parallel wavenet: Fast high-fidelity speech synthesis
Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G., Lockhart, E., Cobo, L., Stimberg, F., et al · 2018
Earlier work this paper cites.
webmushra—a comprehensive framework for web-based listening tests
Schoeffler, M., Bartoschek, S., Stöter, F.-R., Roess, M., Westphal, S., Edler, B., and Herre, J · 2018
Earlier work this paper cites.
Look, listen, and learn more: Design choices for deep audio embeddings
Cramer, A. L., Wu, H.-H., Salamon, J., and Bello, J. P · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Automatic multitrack mixing with a differentiable mixing console of neural audio effects
Steinmetz, C. J., Pons, J., Pascual, S., and Serrà, J · 2020
Earlier work this paper cites.
Neural networks fail to learn periodic functions and how to fix it
Ziyin, L., Hartwig, T., and Ueda, M · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis
Rouard, S. and Hadjeres, G · 2021
Cited alongside, same era.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
Riffusion - stable diffusion for real-time music generation
Forsgren, S. and Martiros, H · 2022
Cited alongside, same era.
Multi-instrument music synthesis with spectrogram diffusion
Hawthorne, C., Simon, I., Roberts, A., Zeghidour, N., Gardner, J., Manilow, E., and Engel, J · 2022
Clipsonic: Text-to-audio synthesis with unlabeled videos and pretrained language-vision models
Dong, H.-W., Liu, X., Pons, J., Bhattacharya, G., Pascual, S., Serrà, J., Berg-Kirkpatrick, T., and McAuley, J · 2023
Later among the works it cites.
Vampnet: Music generation via masked acoustic token modeling
Garcia, H. F., Seetharaman, P., Kumar, R., and Pardo, B · 2023
Later among the works it cites.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
Ghosal, D., Majumder, N., Mehrish, A., and Poria, S · 2023
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2023
Later among the works it cites.
Controllable music production with diffusion models and guidance gradients
Levy, M., Di Giorgi, B., Weers, F., Katharopoulos, A., and Nickson, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mulan: A joint embedding of music audio and natural language
Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P · 2022
Cited alongside, same era.
Efficient training of audio transformers with patchout
Koutini, K., Schlüter, J., Eghbal-zadeh, H., and Widmer, G · 2022
Cited alongside, same era.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Cited alongside, same era.
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Cited alongside, same era.
Novelai improvements on stable diffusion, Oct 2022
NovelAI · 2022
Cited alongside, same era.
Musika! fast infinite waveform music generation
Pasini, M. and Schlüter, J · 2022
Cited alongside, same era.
Li, P., Chen, B., Yao, Y., Wang, Y., Wang, A., and Wang, A · 2023
Later among the works it cites.
Multi-source diffusion models for simultaneous music generation and separation
Mariani, G., Tallini, I., Postolache, E., Mancusi, M., Cosmo, L., and Rodolà, E · 2023
Later among the works it cites.
Solving audio inverse problems with a diffusion model
Moliner, E., Lehtinen, J., and Välimäki, V · 2023
Later among the works it cites.
Full-band general audio synthesis with score-based diffusion
Pascual, S., Bhattacharya, G., Yeh, C., Pons, J., and Serrà, J · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Moûsai: Text-to-music generation with long-context latent diffusion
Schneider, F., Jin, Z., and Schölkopf, B · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts
Vyas, A., Shi, B., Le, M., Tjandra, A., Wu, Y.-C., Guo, B., Zhang, J., Zhang, X., Adkins, R., Ngan, W., et al · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation
Yang, D., Tian, J., Tan, X., Huang, R., Liu, S., Chang, X., Shi, J., Zhao, S., Bian, J., Wu, X., et al · 2023
Later among the works it cites.
Jen-1 composer: A unified framework for high-fidelity multi-track music generation
Yao, Y., Li, P., Chen, B., and Wang, A · 2023
Later among the works it cites.
Efficient neural music generation
Lam, M. W., Tian, Q., Li, T., Yin, Z., Feng, S., Tu, M., Ji, Y., Xia, R., Ma, M., Song, X., et al · 2024
Closest in time.
Common diffusion noise schedules and sample steps are flawed
Lin, S., Liu, B., Li, J., and Yang, X · 2024
Closest in time.
Stemgen: A music generation model that listens
Parker, J., Spijkervet, J., Kosta, K., Yesiler, F., Kuznetsov, B., Wang, J.-C., Avent, M., Chen, J., and Le, D · 2024
Closest in time.
Masked audio generation using a single non-autoregressive transformer
Ziv, A., Gat, I., Lan, G. L., Remez, T., Kreuk, F., Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2024
Closest in time.