Fetching the paper…
Reading the bibliography…
Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music tagging, instrument classification, key detection, and more.
RoBERTa: A robustly optimized Bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Musical genre classification of audio signals
Tzanetakis, G.; and Cook, P. 2002 · 2002
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A.; Qin, J.; Chiu, C.-C.; et al. 2020 · 2005
Earlier work this paper cites.
Evaluation of cnn-based automatic music tagging models
Won, M.; Ferraro, A.; Bogdanov, D.; and Serra, X. 2020 · 2006
Earlier work this paper cites.
The million song dataset
Bertin-Mahieux, T.; Ellis, D. P.; Whitman, B.; and Lamere, P. 2011 · 2011
Earlier work this paper cites.
1000 songs for emotional analysis of music
Soleymani, M.; Caro, M. N.; Schmidt, E. M.; Sha, C.-Y.; and Yang, Y.-H. 2013 · 2013
Earlier work this paper cites.
MIR_EVAL: A transparent implementation of common MIR metrics
Raffel, C.; McFee, B.; Humphrey, E. J.; Salamon, J.; Nieto, O.; Liang, D.; Ellis, D. P.; and Raffel, C. C. 2014 · 2014
Earlier work this paper cites.
Deep Learning and Music Adversaries
Kereliuk, C.; Sturm, B.; and Larsen, J. 2015 · 2015
Earlier work this paper cites.
Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections
Knees, P.; Faraldo Pérez, Á.; Boyer, H.; Vogl, R.; Böck, S.; Hörschläger, F.; Le Goff, M.; et al. 2015 · 2015
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
Engel, J.; Resnick, C.; Roberts, A.; Dieleman, S.; Norouzi, M.; Eck, D.; and Simonyan, K. 2017 · 2017
Earlier work this paper cites.
End-to-End Musical Key Estimation Using a Convolutional Neural Network
Korzeniowski, F.; and Widmer, G. 2017 · 2017
Earlier work this paper cites.
Virtual class enhanced discriminative embedding learning
Chen, B.; Deng, W.; and Shen, H. 2018 · 2018
Earlier work this paper cites.
VocalSet: A Singing Voice Dataset
Wilkins, J.; Seetharaman, P.; Wahl, A.; and Pardo, B. 2018 · 2018
Earlier work this paper cites.
The harmonix set: Beats, downbeats, and functional segment annotations of western popular music
Nieto, O.; McCallum, M. C.; Davies, M. E.; Robertson, A.; Stark, A. M.; and Egozy, E. 2019 · 2019
Earlier work this paper cites.
Fairseq: A fast, extensible toolkit for sequence modeling
Ott, M.; Edunov, S.; Baevski, A.; Fan, A.; Gross, S.; Ng, N.; Grangier, D.; and Auli, M. 2019 · 2019
Cited alongside, same era.
Music4all: A new music database and its applications
Santana, I. A. P.; Pinhelli, F.; Donini, J.; et al. 2020 · 2020
Cited alongside, same era.
Codified audio language modeling learns useful representations for music information retrieval
Castellon, R.; Donahue, C.; and Liang, P. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021 · 2021
Cited alongside, same era.
Contrastive learning of musical representations
Spijkervet, J.; and Burgoyne, J. A. 2021 · 2021
Cited alongside, same era.
Deformable cnn and imbalance-aware feature learning for singing technique classification
Yamamoto, Y.; Nam, J.; and Terasawa, H. 2022 · 2022
Later among the works it cites.
Decoupled contrastive learning
Yeh, C.-H.; Hong, C.-Y.; Hsu, Y.-C.; Liu, T.-L.; Chen, Y.; and LeCun, Y. 2022 · 2022
Later among the works it cites.
MusicLM: Generating music from text
Agostinelli, A.; Denk, T. I.; Borsos, Z.; Engel, J.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; Roberts, A.; Tagliasacchi, M.; et al. 2023 · 2023
Later among the works it cites.
Efficient self-supervised learning with contextualized target representations for vision, speech and language
Baevski, A.; Babu, A.; Hsu, W.-N.; and Auli, M. 2023 · 2023
Later among the works it cites.
Natural Language Supervision for General-Purpose Audio Representations
Elizalde, B.; Deshmukh, S.; and Wang, H. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Supervised chorus detection for popular music using convolutional neural network and multi-task learning
Wang, J.-C.; Smith, J. B.; Chen, J.; Song, X.; and Wang, Y. 2021 · 2021
Cited alongside, same era.
Self-supervised learning with random-projection quantizer for speech recognition
Chiu, C.-C.; Qin, J.; Zhang, Y.; Yu, J.; and Wu, Y. 2022 · 2022
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A.; Copet, J.; Synnaeve, G.; and Adi, Y. 2022 · 2022
Cited alongside, same era.
MuLan: A joint embedding of music audio and natural language
Huang, Q.; Jansen, A.; Lee, J.; Ganti, R.; Li, J. Y.; and Ellis, D. P. W. 2022 · 2022
Cited alongside, same era.
Li, Y.; Yuan, R.; Zhang, G.; Ma, Y.; Lin, C.; Chen, X.; Ragni, A.; Yin, H.; Hu, Z.; He, H.; et al. 2022 · 2022
Cited alongside, same era.
MT4SSL: Boosting self-supervised speech representation learning by integrating multiple targets
Ma, Z.; Zheng, Z.; Tang, C.; Wang, Y.; and Chen, X. 2022 · 2022
Cited alongside, same era.
Self-supervised speech representation learning: A review
Mohamed, A.; Lee, H.-y.; Borgholt, L.; Havtorn, J. D.; Edin, J.; Igel, C.; Kirchhoff, K.; Li, S.-W.; Livescu, K.; Maaløe, L.; et al. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
MERT: Acoustic music understanding model with large-scale self-supervised training
Li, Y.; Yuan, R.; Zhang, G.; et al. 2023 · 2023
Later among the works it cites.
A foundation model for music informatics
Won, M.; Hung, Y.-N.; and Le, D. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Later among the works it cites.
MARBLE: Music audio representation benchmark for universal evaluation
Yuan, R.; Ma, Y.; Li, Y.; et al. 2023 · 2023
Later among the works it cites.
Seed-ASR: Understanding diverse speech and contexts with LLM-based speech recognition
Bai, Y.; Chen, J.; Chen, J.; et al. 2024 · 2024
Later among the works it cites.
EAT: Self-supervised pre-training with efficient audio transformer
Chen, W.; Liang, Y.; Ma, Z.; Zheng, Z.; and Chen, X. 2024 · 2024
Later among the works it cites.
MusiLingo: Bridging music and text with pre-trained language models for music captioning and query response
Deng, Z.; Ma, Y.; Liu, Y.; et al. 2024 · 2024
Later among the works it cites.
Vasilakis, Y.; Bittner, R.; and Pauwels, J. 2024 · 2024
Later among the works it cites.