Fetching the paper…
Reading the bibliography…
In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key, pitch, vocal characteristics, tempo and sonority.
G. Tzanetakis and P. Cook, “Musical genre classification of audio signals,” IEEE Transactions on Speech and Audio Processing , vol. 10, no. 5, pp. 293–302, 2002
2002
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging.” in ISMIR , 2009
2009
Earlier work this paper cites.
M. Soleymani, M. N. Caro, E. M. Schmidt, C.-Y. Sha, and Y.-H. Yang, “1000 songs for emotional analysis of music,” in CrowdMM , 2013, pp. 1–6
2013
Earlier work this paper cites.
F. Weninger, F. Eyben, and B. Schuller, “On-line continuous-time music mood regression with deep recurrent neural networks,” in ICASSP , 2014, pp. 5412–5416
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” ICLR , 2015
2015
Earlier work this paper cites.
P. Knees, A. Faraldo, P. Herrera, R. Vogl, S. Böck, F. Hörschläger, and M. Le Goff, “Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections,” in ISMIR , 2015
2015
Earlier work this paper cites.
Pioneer DJ, https://rekordbox.com , [Accessed: 2016-09-12]
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017, pp. 776–780
2017
Earlier work this paper cites.
F. Korzeniowski and G. Widmer, “End-to-end musical key estimation using a convolutional neural network,” in EUSIPCO , 2017, pp. 966–970
2017
Earlier work this paper cites.
G. Bernardes, M. E. Davies, and C. Guedes, “Automatic musical key estimation with adaptive mode bias,” in ICASSP , 2017, pp. 316–320
2017
Earlier work this paper cites.
J. Pons, O. Nieto, M. Prockup, E. Schmidt, A. Ehmann, and X. Serra, “End-to-end learning for music audio tagging at scale,” ISMIR , 2018
2018
Earlier work this paper cites.
A. Jansen, M. Plakal, R. Pandya, D. P. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, “Unsupervised learning of semantic audio representations,” in ICASSP , 2018, pp. 126–130
2018
Earlier work this paper cites.
J. Lee, J. Park, K. L. Kim, and J. Nam, “SampleCNN: End-to-end deep convolutional neural networks using very small filters for music classification,” Applied Sciences , vol. 8, no. 1, p. 150, 2018
2018
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” ICLR , 2018
2018
Earlier work this paper cites.
S. Oramas, F. Barbieri, O. Nieto Caballero, and X. Serra, “Multimodal deep learning for music genre classification,” Transactions of the International Society for Music Information Retrieval , vol. 1, no. 1, pp. 4–21, 2018
2018
Earlier work this paper cites.
J. Yan, Y. Song, W. Guo, L.-R. Dai, I. McLoughlin, and L. Chen, “A region based attention method for weakly supervised sound event detection and classification,” in ICASSP , 2019, pp. 755–759
2019
Earlier work this paper cites.
M. C. McCallum, “Unsupervised learning of deep features for music segmentation,” in ICASSP , 2019, pp. 346–350
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra, “The MTG-Jamendo dataset for automatic music tagging,” in ICML Workshops , 2019. [Online]. Available: http://hdl.handle.net/10230/42015
2019
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
M. Tagliasacchi, B. Gfeller, F. de Chaumont Quitry, and D. Roblek, “Pre-training audio representations with self-supervision,” IEEE Signal Processing Letters , vol. 27, pp. 600–604, 2020
2020
Cited alongside, same era.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive learning of general-purpose audio representations,” in ICASSP , 2021, pp. 3875–3879
2021
Later among the works it cites.
J. Spijkervet and J. A. Burgoyne, “Contrastive learning of musical representations,” ISMIR , 2021
2021
Later among the works it cites.
Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio spectrogram transformer,” Interspeech , 2021
2021
Later among the works it cites.
W.-T. Lu, J.-C. Wang, M. Won, K. Choi, and X. Song, “SpecTNT: a time-frequency transformer for music audio,” ISMIR , 2021
2021
Later among the works it cites.
M. Won, K. Choi, and X. Serra, “Semi-supervised music tagging transformer,” ISMIR , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Wang and A. Oord, “Multi-format contrastive learning of audio representations,” NeurIPS Workshops , 2020
2020
Cited alongside, same era.
Q. Huang, A. Jansen, L. Zhang, D. P. Ellis, R. A. Saurous, and J. Anderson, “Large-scale weakly-supervised content embeddings for music recommendation and tagging,” in ICASSP , 2020, pp. 8364–8368
2020
Cited alongside, same era.
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” in ICASSP , 2020, pp. 6419–6423
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Kim, J. Urbano, C. Liem, and A. Hanjalic, “One deep music representation to rule them all? a comparative analysis of different representation learning strategies,” Neural Computing and Applications , vol. 32, no. 4, pp. 1067–1093, 2020
2020
Cited alongside, same era.
T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. E. Hinton, “Big self-supervised models are strong semi-supervised learners,” Advances in neural information processing systems , vol. 33, pp. 22 243–22 255, 2020
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML , 2020, pp. 1597–1607
2020
Cited alongside, same era.
Y. Gong, Y.-A. Chung, and J. Glass, “PSLA: Improving audio tagging with pretraining, sampling, labeling, and aggregation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3292–3306, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Won, S. Oramas, O. Nieto, F. Gouyon, and X. Serra, “Multimodal metric learning for tag-based music retrieval,” in ICASSP , 2021, pp. 591–595
2021
Later among the works it cites.
R. Castellon, C. Donahue, and P. Liang, “Codified audio language modeling learns useful representations for music information retrieval,” ISMIR , 2021
2021
Later among the works it cites.
F. Korzeniowski, O. Nieto, M. McCallum, M. Won, S. Oramas, and E. Schmidt, “Mood classification using listening data,” in ISMIR , 2021
2021
Later among the works it cites.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for audio: Self-supervised learning for general-purpose audio representation,” in IJCNN , 2021, pp. 1–8
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Wang, P. Luc, Y. Wu, A. Recasens, L. Smaira, A. Brock, A. Jaegle, J.-B. Alayrac, S. Dieleman, J. Carreira et al. , “Towards learning universal audio representations,” in ICASSP , 2022, pp. 4593–4597
2022
Closest in time.
I. Manco, E. Benetos, E. Quinton, and G. Fazekas, “Learning music audio representations via weak language supervision,” in ICASSP , 2022, pp. 456–460
2022
Closest in time.
HEAR Benchmark, https://hearbenchmark.com/ , [Accessed: 2022-08-01]
2022
Closest in time.