Fetching the paper…
Reading the bibliography…
Automatic music captioning, which generates natural language descriptions for given music tracks, holds significant potential for enhancing the understanding and organization of large volumes of musical data.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics (ACL) , 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005
2005
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging.” in International Conference on Music Information Retrieval (ISMIR) , 2009
2009
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The million song dataset,” in International Conference on Music Information Retrieval (ISMIR) , 2011
2011
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
K. Choi, G. Fazekas, B. McFee, K. Cho, and M. Sandler, “Towards music captioning: Generating music playlist descriptions,” in International Society for Music Information Retrieval Conference (ISMIR), Late-Breaking/Demo , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Proceedings of the Advances in neural information processing systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
K. Choi, G. Fazekas, K. Cho, and M. Sandler, “The effects of noisy labels on deep convolutional neural networks for music tagging,” IEEE Transactions on Emerging Topics in Computational Intelligence , 2018
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019
2019
Earlier work this paper cites.
T. Cai, M. I. Mandel, and D. He, “Music autotagging as captioning,” in Proceedings of the 1st Workshop on NLP for Music and Audio (NLP4MusA) , 2020
2020
Earlier work this paper cites.
M. Won, S. Chun, O. Nieto, and X. Serrc, “Data-driven harmonic filters for audio representation learning,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in Proceedings of the Advances in neural information processing systems (NeurIPS) , 2020
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, jan 2020
2020
Cited alongside, same era.
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy, “The pile: An 800gb dataset of diverse text for language modeling,” 2020
2020
Cited alongside, same era.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “Bertscore: Evaluating text generation with bert,” in International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” in Proceedings of the Advances in neural information processing systems (NeurIPS) , 2022
2022
Later among the works it cites.
Y. Zhang, J. Jiang, G. Xia, and S. Dixon, “Interpreting song lyrics with an audio-informed pre-trained language model,” in International Conference on Music Information Retrieval (ISMIR) , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Stefanini, M. Cornia, L. Baraldi, S. Cascianelli, G. Fiameni, and R. Cucchiara, “From show to tell: a survey on deep learning-based image captioning,” IEEE transactions on pattern analysis and machine intelligence , 2022
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) , 2020
2020
Cited alongside, same era.
I. Manco, E. Benetos, E. Quinton, and G. Fazekas, “Muscaps: Generating captions for music audio,” in International Joint Conference on Neural Networks (IJCNN) . IEEE, 2021
2021
Cited alongside, same era.
S. Doh, J. Lee, and J. Nam, “Music playlist title generation: A machine-translation approach,” in Proceedings of the 2nd Workshop on NLP for Music and Spoken Audio (NLP4MuSA) , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
G. Gabbolini, R. Hennequin, and E. Epure, “Data-efficient playlist captioning with musical and linguistic knowledge,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2022
2022
Cited alongside, same era.
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. Ellis, “MuLan: A joint embedding of music audio and natural language,” in International Conference on Music Information Retrieval (ISMIR) , 2022
2022
Cited alongside, same era.
I. Manco, B. Weck, P. Tovstogan, M. Won, and D. Bogdanov, “Song describer: a platform for collecting textual descriptions of music recordings,” in International Conference on Music Information Retrieval (ISMIR), Late-Breaking/Demo session , 2022
2022
Cited alongside, same era.
T. Chen, Y. Xie, S. Zhang, S. Huang, H. Zhou, and J. Li, “Learning music sequence representation from text supervision,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022
2022
Cited alongside, same era.
Later among the works it cites.
H. Kim, S. Doh, J. Lee, and J. Nam, “Music playlist title generation using artist information,” in Proceedings of the AAAI-23 Workshop on Creative AI Across Modalities , 2023
2023
Closest in time.
2023
Closest in time.
S. Doh, M. Won, K. Choi, and J. Nam, “Toward universal text-to-music retrieval,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys , 2023
2023
Closest in time.