Fetching the paper…
Reading the bibliography…
Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio.
Automatic musical genre classification of audio signals
George, T., Georg, E., and Perry, C · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A · 2005
Earlier work this paper cites.
Evaluation of algorithms using games: The case of music tagging
Law, E., West, K., Mandel, M., Bay, M., and Downie, J. S · 2009
Earlier work this paper cites.
Classification accuracy is not enough: On the evaluation of music genre recognition systems
Sturm, B. L · 2013
Earlier work this paper cites.
MedleyDB: A multitrack dataset for annotation-intensive MIR research
Bittner, R. M., Salamon, J., Tierney, M., Mauch, M., Cannam, C., and Bello, J. P · 2014
Earlier work this paper cites.
From machine learning to machine reasoning: An essay
Bottou, L · 2014
Earlier work this paper cites.
mir_eval: A transparent implementation of common mir metrics
Raffel, C., McFee, B., Humphrey, E. J., Salamon, J., Nieto, O., Liang, D., Ellis, D. P., and Raffel, C. C · 2014
Earlier work this paper cites.
Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections
Knees, P., Faraldo Pérez, Á., Boyer, H., Vogl, R., Böck, S., Hörschläger, F., Le Goff, M., et al · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S · 2016
Earlier work this paper cites.
madmom: a new Python Audio and Music Signal Processing Library
Böck, S., Korzeniowski, F., Schlüter, J., Krebs, F., and Widmer, G · 2016
Earlier work this paper cites.
Key estimation in electronic dance music
Faraldo, Á., Gómez, E., Jordà, S., and Herrera, P · 2016
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R. and Zisserman, A · 2017
Earlier work this paper cites.
Fma: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
End-to-end musical key estimation using a convolutional neural network
Korzeniowski, F. and Widmer, G · 2017
Earlier work this paper cites.
Learning features of music from scratch
Thickstun, J., Harchaoui, Z., and Kakade, S · 2017
Earlier work this paper cites.
Automatic music transcription: An overview
Benetos, E., Dixon, S., Duan, Z., and Ewert, S · 2018
Earlier work this paper cites.
Musical source separation: An introduction
Cano, E., FitzGerald, D., Liutkus, A., Plumbley, M. D., and Stöter, F.-R · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Hudson, D. A. and Manning, C. D · 2018
Earlier work this paper cites.
Frame-level instrument recognition by timbre and pitch
Hung, Y.-N. and Yang, Y.-H · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Bogdanov, D., Won, M., Tovstogan, P., Porter, A., and Serra, X · 2019
Earlier work this paper cites.
Multitask learning for frame-level instrument recognition
Hung, Y.-N., Chen, Y.-A., and Yang, Y.-H · 2019
Cited alongside, same era.
Cutting music source separation some slakh: A dataset to study the impact of training data quality and quantity
Manilow, E., Wichern, G., Seetharaman, P., and Le Roux, J · 2019
Cited alongside, same era.
Model cards for model reporting
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T · 2019
Cited alongside, same era.
20 years of automatic chord recognition from audio
Pauwels, J., O’Hanlon, K., Gómez, E., Sandler, M., et al · 2019
Cited alongside, same era.
Musical tempo and key estimation using convolutional neural networks with directional filters
Schreiber, H. and Müller, M · 2019
Cited alongside, same era.
Wav2clip: Learning robust audio representations from clip
Wu, H.-H., Seetharaman, P., Kumar, K., and Bello, J. P · 2022
Later among the works it cites.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Closest in time.
Pali: A jointly-scaled multilingual language-image model
Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S. A., Grycner, A., Mustafa, B., Beyer, L., Kolesnikov, A., Puigcerver, J., Ding, N., Rong, K., Akbari, H., Mishra, G., Xue, L., Thapliyal, A., Bradbury, J., Kuo, W., Seyedhosseini, M., Jia, C., Ayan, B. K., Riquelme, C., Steiner, A., Angelova, A., Zhai, X., Houlsby, N., and Soricut, R · 2023
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simonetta, F., Ntalampiras, S., and Avanzini, F · 2019
Cited alongside, same era.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Cited alongside, same era.
The freesound loop dataset and annotation tool
Ramires, A., Font, F., Bogdanov, D., Smith, J. B., Yang, Y.-H., Ching, J., Chen, B.-Y., Wu, Y.-K., Wei-Han, H., and Serra, X · 2020
Cited alongside, same era.
Music tempo estimation: Are we done yet?
Schreiber, H., Urbano, J., and Müller, M · 2020
Cited alongside, same era.
Codified audio language modeling learns useful representations for music information retrieval
Castellon, R., Donahue, C., and Liang, P · 2021
Cited alongside, same era.
Music tempo estimation via neural networks–a comparative analysis
de Souza, M. S. d. O., Moura, P. N. d. S., and Briot, J.-P · 2021
Cited alongside, same era.
Mt3: Multi-task multitrack music transcription
Gardner, J. P., Simon, I., Manilow, E., Hawthorne, C., and Engel, J · 2021
Cited alongside, same era.
Deshmukh, S., Elizalde, B., Singh, R., and Wang, H · 2023
Closest in time.
Lp-musiccaps: Llm-based pseudo music captioning
Doh, S., Choi, K., Lee, J., and Nam, J · 2023
Closest in time.
Clap learning audio concepts from natural language supervision
Elizalde, B., Deshmukh, S., Al Ismail, M., and Wang, H · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I · 2023
Closest in time.
Imagebind-llm: Multi-modality instruction tuning
Han, J., Zhang, R., Shao, W., Gao, P., Xu, P., Xiao, H., Zhang, K., Liu, C., Wen, S., Guo, Z., et al · 2023
Closest in time.
A whisper transformer for audio captioning trained with synthetic captions and transfer learning
Kadlčík, M., Hájek, A., Kieslich, J., and Winiecki, R · 2023
Closest in time.
Mert: Acoustic music understanding model with large-scale self-supervised training
Li, Y., Yuan, R., Zhang, G., Ma, Y., Chen, X., Yin, H., Lin, C., Ragni, A., Benetos, E., Gyenge, N., et al · 2023
Closest in time.
Language-guided music recommendation for video via prompt analogies
McKee, D., Salamon, J., Sivic, J., and Russell, B · 2023
Closest in time.
Anymal: An efficient and scalable any-modality augmented language model
Moon, S., Madotto, A., Lin, Z., Nagarajan, T., Smith, M., Jain, S., Yeh, C.-F., Murugesan, P., Heidari, P., Liu, Y., et al · 2023
Closest in time.
Improving multimodal datasets with image captioning
Nguyen, T., Gadre, S. Y., Ilharco, G., Oh, S., and Schmidt, L · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Deep learning methods for instrument separation and recognition
Vianna Lordelo, C · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Closest in time.
Llm evaluators recognize and favor their own generations
Panickssery, A., Bowman, S. R., and Feng, S · 2024
Closest in time.