Fetching the paper…
Reading the bibliography…
Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models.
“Training deep nets with sublinear memory cost,”
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin, · 2016
Earlier work this paper cites.
“AudioSet: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Earlier work this paper cites.
“Fréchet Audio Distance: A metric for evaluating music enhancement algorithms,”
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi, · 2018
Earlier work this paper cites.
“Look, listen, and learn more: Design choices for deep audio embeddings,”
Aurora Linh Cramer, Ho-Hsiang Wu, Justin Salamon, and Juan Pablo Bello, · 2019
Earlier work this paper cites.
“AudioCaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer,”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu, · 2020
Earlier work this paper cites.
“Neural networks fail to learn periodic functions and how to fix it,”
Liu Ziyin, Tilman Hartwig, and Masahito Ueda, · 2020
Earlier work this paper cites.
“Jukebox: A generative model for music,”
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever, · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley, · 2020
Earlier work this paper cites.
“auraloss: Audio focused loss functions in PyTorch,”
Christian J. Steinmetz and Joshua D. Reiss, · 2020
Earlier work this paper cites.
“Audiogen: Textually guided audio generation,”
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, · 2022
Cited alongside, same era.
“FlashAttention: Fast and memory-efficient exact attention with IO-awareness,”
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Progressive distillation for fast sampling of diffusion models,”
Tim Salimans and Jonathan Ho, · 2022
Cited alongside, same era.
“DPM-solver++: Fast solver for guided sampling of diffusion probabilistic models,”
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu, · 2022
Cited alongside, same era.
“Efficient training of audio transformers with patchout,”
“Scalable diffusion models with transformers,”
William Peebles and Saining Xie, · 2023
Later among the works it cites.
“Controllable music production with diffusion models and guidance gradients,”
Mark Levy, Bruno Di Giorgi, Floris Weers, Angelos Katharopoulos, and Tom Nickson, · 2023
Later among the works it cites.
“RoFormer: Enhanced transformer with rotary position embedding,”
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu, · 2023
Later among the works it cites.
“Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,”
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao, · 2023
Later among the works it cites.
“The Song Describer Dataset: a corpus of audio captions for music-and-language evaluation,”
Ilaria Manco, Benno Weck, SeungHeon Doh, Minz Won, Yixiao Zhang, Dmitry Bogdanov, Yusong Wu, Ke Chen, Philip Tovstogan, Emmanouil Benetos, Elio Quinton, György Fazekas, and Juhan Nam, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2022
Cited alongside, same era.
“MusicLM: Generating music from text,”
Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank, · 2023
Cited alongside, same era.
“AudioLDM: Text-to-audio generation with latent diffusion models,”
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley, · 2023
Cited alongside, same era.
“AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,”
Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley, · 2023
Cited alongside, same era.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2023
Cited alongside, same era.
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2023
Cited alongside, same era.
“Audiogenai/agc: Audiogen codec,”
AudiogenAI,
Cited in the paper.
“Extracting training data from diffusion models,”
Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace, · 2023
Later among the works it cites.
“Fast timing-conditioned latent audio diffusion,”
Zach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley, and Jordi Pons, · 2024
Closest in time.
“Long-form music generation with latent diffusion,”
Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons, · 2024
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2024
Closest in time.
“Scaling rectified flow transformers for high-resolution image synthesis,”
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al., · 2024
Closest in time.