Fetching the paper…
Reading the bibliography…
We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models.
BLEU : a Method for Automatic Evaluation of Machine Translation
K. Papineni, S. Roukos, T. Ward, and W.-j. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C. Y. Lin · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
A. Lavie and A. Agarwal · 2007
Earlier work this paper cites.
The million song dataset
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere · 2011
Earlier work this paper cites.
The Musicality of Non-Musicians: An Index for Assessing Musical Sophistication in the General Population
D. Müllensiefen, B. Gingras, J. Musil, and L. Stewart · 2014
Earlier work this paper cites.
YouTube-8M: A Large-Scale Video Classification Benchmark, Sept. 2016
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
DeLiGAN: Generative Adversarial Networks for Diverse and Limited Data
S. Gurumurthy, R. K. Sarvadevabhatla, and R. V. Babu · 2017
Earlier work this paper cites.
The MTG-Jamendo Dataset for Automatic Music Tagging
D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra · 2019
Earlier work this paper cites.
Fréchet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi · 2019
Earlier work this paper cites.
Do ImageNet Classifiers Generalize to ImageNet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2020
Earlier work this paper cites.
What Will it Take to Fix Benchmarking in Natural Language Understanding?, Oct. 2021
S. R. Bowman and G. E. Dahl · 2021
Cited alongside, same era.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. III, and K. Crawford · 2021
Cited alongside, same era.
MusCaps: Generating Captions for Music Audio
I. Manco, E. Benetos, E. Quinton, and G. Fazekas · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Toward Universal Text-to-Music Retrieval
S. Doh, M. Won, K. Choi, and J. Nam · 2022
Cited alongside, same era.
MusicLM: Generating Music From Text, Jan. 2023
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank · 2023
Closest in time.
Simple and Controllable Music Generation, June 2023
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez · 2023
Closest in time.
LP-MusicCaps: LLM-Based Pseudo Music Captioning
S. Doh, K. Choi, J. Lee, and J. Nam · 2023
Closest in time.
Noise2Music: Text-conditioned Music Generation with Diffusion Models, Mar. 2023
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank, J. Engel, Q. V. Le, W. Chan, Z. Chen, and W. Han · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-Efficient Playlist Captioning With Musical and Linguistic Knowledge
G. Gabbolini, R. Hennequin, and E. Epure · 2022
Cited alongside, same era.
MuLan: A Joint Embedding of Music Audio and Natural Language
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis · 2022
Cited alongside, same era.
Conversational Music Retrieval with Synthetic Data
M. E. Leszczynski, R. Ganti, S. Zhang, K. Balog, F. Radlinski, F. Pereira, and A. T. Chaganty · 2022
Cited alongside, same era.
Contrastive Audio-Language Learning for Music
I. Manco, E. Benetos, E. Quinton, and G. Fazekas · 2022
Cited alongside, same era.
Learning music audio representations via weak language supervision
I. Manco, E. Benetos, E. Quinton, and G. Fazekas · 2022
Cited alongside, same era.
Song Describer: a Platform for Collecting Textual Descriptions of Music Recordings
I. Manco, B. Weck, P. Tovstogan, M. Won, and D. Bogdanov · 2022
Cited alongside, same era.
Riffusion - Stable diffusion for real-time music generation, 2022
H. M. Seth Forsgren · 2022
Cited alongside, same era.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley · 2023
Closest in time.
S. Liu, A. S. Hussain, C. Sun, and Y. Shan · 2023
Closest in time.
Language-Guided Music Recommendation for Video via Prompt Analogies, June 2023
D. McKee, J. Salamon, J. Sivic, and B. Russell · 2023
Closest in time.
Mousai: Text-to-Music Generation with Long-Context Latent Diffusion, Jan. 2023
F. Schneider, Z. Jin, and B. Schölkopf · 2023
Closest in time.
When does dough become a bagel?Analyzing the remaining mistakes on ImageNet
V. Vasudevan, B. Caine, R. Gontijo-Lopes, S. Fridovich-Keil, and R. Roelofs · 2023
Closest in time.
Data Leakage in Cross-Modal Retrieval Training: A Case Study
B. Weck and X. Serra · 2023
Closest in time.
Audio-Text Models Do Not Yet Leverage Natural Language, Mar. 2023
H.-H. Wu, O. Nieto, J. P. Bello, and J. Salamon · 2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Closest in time.