Fetching the paper…
Reading the bibliography…
In the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding.
musicnn: Pre-trained convolutional neural networks for music audio tagging
Pons, J. and Serra, X. (2019) · 1909
Earlier work this paper cites.
Polyscriber: Integrated fine-tuning of extractor and lyrics transcriber for polyphonic music
Gao, X., Gupta, C., and Li, H. (2023) · 1981
Earlier work this paper cites.
Musical genre classification of audio signals
Tzanetakis, G. and Cook, P. (2002) · 2002
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. (2020) · 2004
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I. (2020) · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F. J., and Schmidhuber, J. (2006) · 2006
Earlier work this paper cites.
Evaluation of algorithms using games: The case of music tagging
Law, E., West, K., Mandel, M. I., Bay, M., and Downie, J. S. (2009) · 2009
Earlier work this paper cites.
The million song dataset
Bertin-Mahieux, T., Ellis, D. P., Whitman, B., and Lamere, P. (2011) · 2011
Earlier work this paper cites.
1000 songs for emotional analysis of music
Soleymani, M., Caro, M. N., Schmidt, E. M., Sha, C.-Y., and Yang, Y.-H. (2013) · 2013
Earlier work this paper cites.
Medleydb: A multitrack dataset for annotation-intensive mir research
Bittner, R., Salamon, J., Tierney, M., Mauch, M., Cannam, C., and Bello, J. (2014) · 2014
Earlier work this paper cites.
Mir_eval: A transparent implementation of common mir metrics
Raffel, C., McFee, B., Humphrey, E. J., Salamon, J., Nieto, O., Liang, D., Ellis, D. P., and Raffel, C. C. (2014) · 2014
Earlier work this paper cites.
Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections
Knees, P., Faraldo Pérez, Á., Boyer, H., Vogl, R., Böck, S., Hörschläger, F., Le Goff, M., et al. (2015) · 2015
Earlier work this paper cites.
Swing ratio estimation
Marchand, U. and Peeters, G. (2015) · 2015
Earlier work this paper cites.
Fundamentals of music processing: Audio, analysis, algorithms, applications
Müller, M. (2015) · 2015
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
Engel, J., Resnick, C., Roberts, A., Dieleman, S., Norouzi, M., Eck, D., and Simonyan, K. (2017) · 2017
Earlier work this paper cites.
End-to-end musical key estimation using a convolutional neural network
Korzeniowski, F. and Widmer, G. (2017) · 2017
Earlier work this paper cites.
The MUSDB18 corpus for music separation
Rafii, Z., Liutkus, A., Stöter, F.-R., Mimilakis, S. I., and Bittner, R. (2017) · 2017
Earlier work this paper cites.
Hybrid ctc/attention architecture for end-to-end speech recognition
Watanabe, S., Hori, T., Kim, S., Hershey, J. R., and Hayashi, T. (2017) · 2017
Earlier work this paper cites.
Music style transfer: A position paper
Dai, S., Zhang, Z., and Xia, G. G. (2018) · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Earlier work this paper cites.
Vocalset: A singing voice dataset
Wilkins, J., Seetharaman, P., Wahl, A., and Pardo, B. (2018) · 2018
Earlier work this paper cites.
Guitarset: A dataset for guitar transcription
Xi, Q., Bittner, R. M., Pauwels, J., Ye, X., and Bello, J. P. (2018) · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Bogdanov, D., Won, M., Tovstogan, P., Porter, A., and Serra, X. (2019) · 2019
Cited alongside, same era.
Scaling and benchmarking self-supervised visual representation learning
Goyal, P., Mahajan, D., Gupta, A., and Misra, I. (2019) · 2019
Cited alongside, same era.
Large-vocabulary chord transcription via chord structure decomposition
Jiang, J., Chen, K., Li, W., and Xia, G. (2019) · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K. (2019) · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. (2019) · 2019
Cited alongside, same era.
The visual task adaptation benchmark
Multi-task self-supervised pre-training for music classification
Wu, H.-H., Kao, C.-C., Tang, Q., Sun, M., McFee, B., Bello, J. P., and Wang, C. (2021) · 2021
Later among the works it cites.
Superb: Speech processing universal performance benchmark
Yang, S.-w., Chi, P.-H., Chuang, Y.-S., Lai, C.-I. J., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G.-T., et al. (2021) · 2021
Later among the works it cites.
Musicoder: A universal music-acoustic encoder based on transformer
Zhao, Y. and Guo, J. (2021) · 2021
Later among the works it cites.
Music representation learning based on editorial metadata from discogs
Alonso-Jiménez, P., Serra, X., and Bogdanov, D. (2022) · 2022
Later among the works it cites.
Mulan: A joint embedding of music audio and natural language
Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al. (2019) · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. (2020) · 2020
Cited alongside, same era.
Eraser: A benchmark to evaluate rationalized nlp models
DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., and Wallace, B. C. (2020) · 2020
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D. (2020) · 2020
Cited alongside, same era.
Deep contextualized acoustic representations for semi-supervised speech recognition
Ling, S., Liu, Y., Salazar, J., and Kirchhoff, K. (2020) · 2020
Cited alongside, same era.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Liu, A. T., Yang, S.-w., Chi, P.-H., Hsu, P.-c., and Lee, H.-y. (2020) · 2020
Cited alongside, same era.
How useful is self-supervised pretraining for visual tasks?
Newell, A. and Deng, J. (2020) · 2020
Cited alongside, same era.
Map-music2vec: A simple and effective baseline for self-supervised music audio representation learning
Li, Y., Yuan, R., Zhang, G., MA, Y., Lin, C., Chen, X., Ragni, A., Yin, H., Hu, Z., He, H., et al. (2022b) · 2022
Later among the works it cites.
Learning music audio representations via weak language supervision
Manco, I., Benetos, E., Quinton, E., and Fazekas, G. (2022) · 2022
Later among the works it cites.
Supervised and unsupervised learning of audio representations for music understanding
McCallum, M. C., Korzeniowski, F., Oramas, S., Gouyon, F., and Ehmann, A. F. (2022) · 2022
Later among the works it cites.
Transfer learning with deep neural embeddings for music classification tasks
Modrzejewski, M., Szachewicz, P., and Rokita, P. (2023) · 2022
Later among the works it cites.
Transfer learning of wav2vec 2.0 for automatic lyric transcription
Ou, L., Gu, X., and Wang, Y. (2022) · 2022
Later among the works it cites.
Towards learning universal audio representations
Wang, L., Luc, P., Wu, Y., Recasens, A., Smaira, L., Brock, A., Jaegle, A., Alayrac, J.-B., Dieleman, S., Carreira, J., et al. (2022a) · 2022
Later among the works it cites.
Deformable cnn and imbalance-aware feature learning for singing technique classification
Yamamoto, Y., Nam, J., and Terasawa, H. (2022) · 2022
Later among the works it cites.
Contrastive learning with positive-negative frame mask for music representation
Yao, D., Zhao, Z., Zhang, S., Zhu, J., Zhu, Y., Zhang, R., and He, X. (2022) · 2022
Later among the works it cites.
Transfer learning with jukebox for music source separation
Zai El Amri, W., Tautz, O., Ritter, H., and Melnik, A. (2022) · 2022
Later among the works it cites.
Contrastive learning-based audio to lyrics alignment for multiple languages
Durand, S., Stoller, D., and Ewert, S. (2023a) · 2023
Closest in time.
Contrastive learning-based audio to lyrics alignment for multiple languages
Durand, S., Stoller, D., and Ewert, S. (2023b) · 2023
Closest in time.
Mert: Acoustic music understanding model with large-scale self-supervised training
Li, Y., Yuan, R., Zhang, G., Ma, Y., Chen, X., Yin, H., Lin, C., Ragni, A., Benetos, E., Gyenge, N., Dannenberg, R., Liu, R., Chen, W., Xia, G., Shi, Y., Huang, W., Guo, Y., and Fu, J. (2023) · 2023
Closest in time.
On the effectiveness of speech self-supervised learning for music
Ma, Y., Yuan, R., Li, Y., Zhang, G., Chen, X., Yin, H., Lin, C., Benetos, E., Ragni, A., Gyenge, N., et al. (2023) · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. (2023) · 2023
Closest in time.
Hybrid transformers for music source separation
Rouard, S., Massa, F., and Défossez, A. (2023) · 2023
Closest in time.
Lyricwhiz: Robust multilingual lyrics transcription by whispering to chatgpt
Zhuo, L., Yuan, R., Pan, J., Ma, Y., Li, Y., Zhang, G., Liu, S., Dannenberg, R., Fu, J., Lin, C., et al. (2023) · 2023
Closest in time.
Deep learning and music adversaries
Kereliuk, C., Sturm, B. L., and Larsen, J. (2015) · 2071
Closest in time.