Fetching the paper…
Reading the bibliography…
Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions.
“BLEU: A Method for Automatic Evaluation of Machine Translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“ROUGE: A Package for Automatic Evaluation of Summaries,”
Chin-Yew Lin, · 2004
Earlier work this paper cites.
“METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,”
Satanjeev Banerjee and Alon Lavie, · 2005
Earlier work this paper cites.
“Evaluation of Algorithms Using Games: The Case of Music Tagging,”
Edith Law, Kris West, Michael I Mandel, et al., · 2009
Earlier work this paper cites.
“Constant-Q Transform Toolbox for Music Processing,”
Christian Schörkhuber and Anssi Klapuri, · 2010
Earlier work this paper cites.
“Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings,”
Aurora Linh Cramer, Ho-Hsiang Wu, Justin Salamon, and Juan Pablo Bello, · 2019
Earlier work this paper cites.
“The MTG-Jamendo Dataset for Automatic Music Tagging,”
Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra, · 2019
Earlier work this paper cites.
“Jukebox: A Generative Model for Music,”
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever, · 2020
Earlier work this paper cites.
“BERTScore: Evaluating Text Generation with BERT,”
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi, · 2020
Earlier work this paper cites.
“MusCaps: Generating Captions for Music Audio,”
Ilaria Manco, Emmanouil Benetos, Elio Quinton, and György Fazekas, · 2021
Earlier work this paper cites.
“Audio Captioning Transformer,”
Xinhao Mei, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang, · 2021
Cited alongside, same era.
“Codified Audio Language Modeling Learns Useful Representations for Music Information Retrieval,”
Rodrigo Castellon, Chris Donahue, and Percy Liang, · 2021
Cited alongside, same era.
“Learning Transferable Visual Models from Natural Language Supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, et al., · 2021
Cited alongside, same era.
“The SJTU System for DCASE2022 Challenge Task 6: Audio Captioning with Audiotext Retrieval Pre-training,”
Xuenan Xu, Zeyu Xie, Mengyue Wu, and Kai Yu, · 2022
Cited alongside, same era.
“Wav2CLIP: Learning Robust Audio Representations From CLIP,”
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
Cited alongside, same era.
“LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention,”
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao, · 2023
Closest in time.
“Unified Model for Image, Video, Audio and Language Tasks,”
Mustafa Shukor, Corentin Dancette, Alexandre Rame, and Matthieu Cord, · 2023
Closest in time.
“Listen, Think, and Understand,”
Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass, · 2023
Closest in time.
“ImageBind: One Embedding Space To Bind Them All,”
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chieh-Hsin Lai, Dongmian Zou, and Gilad Lerman, · 2022
Cited alongside, same era.
“Simple and Controllable Music Generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2023
Cited alongside, same era.
“Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion,”
Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf, · 2023
Cited alongside, same era.
“MusicLM: Generating Music From Text,”
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al., · 2023
Cited alongside, same era.
“LP-MusicCaps: LLM-Based Pseudo Music Captioning,”
SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam, · 2023
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al., · 2023
Closest in time.
“LLaMA 2: Open Foundation and Fine-Tuned Chat Models,”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al., · 2023
Closest in time.
“Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs,” 2023
MosaicML NLP Team et al., · 2023
Closest in time.
“Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,”
Yusong Wu, Ke Chen, Tianyu Zhang, et al., · 2023
Closest in time.
“MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,”
Yizhi Li, Ruibin Yuan, Ge Zhang, et al., · 2023
Closest in time.
“LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model,”
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al., · 2023
Closest in time.