Fetching the paper…
Reading the bibliography…
We demonstrate that language models pre-trained on codified (discretely-encoded) music audio learn representations that are useful for downstream MIR tasks.
G. Tzanetakis and P. Cook, “Musical genre classification of audio signals,” IEEE Transactions on Speech and Audio Processing , 2002
2002
Earlier work this paper cites.
D. Eck and J. Schmidhuber, “Finding temporal structure in music: Blues improvisation with LSTM recurrent networks,” in IEEE Workshop on Neural Networks for Signal Processing , 2002
2002
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging.” in ISMIR , 2009
2009
Earlier work this paper cites.
A. Huq, J. P. Bello, and R. Rowe, “Automated music emotion recognition: A systematic evaluation,” Journal of New Music Research , 2010
2010
Earlier work this paper cites.
P. Hamel and D. Eck, “Learning features from music audio with deep belief networks,” in ISMIR , 2010
2010
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The Million Song Dataset,” in ISMIR , 2011
2011
Earlier work this paper cites.
J. Weston, S. Bengio, and P. Hamel, “Multi-tasking with joint semantic spaces for large-scale music annotation and retrieval,” Journal of New Music Research , 2011
2011
Earlier work this paper cites.
K. Seyerlehner, M. Schedl, R. Sonnleitner, D. Hauger, and B. Ionescu, “From improved auto-taggers to improved music similarity measures,” in International Workshop on Adaptive Multimedia Retrieval , 2012
2012
Earlier work this paper cites.
P. Hamel, M. Davies, K. Yoshii, and M. Goto, “Transfer learning in MIR: Sharing learned latent representations for music audio classification and similarity,” in ISMIR , 2013
2013
Earlier work this paper cites.
M. Soleymani, M. N. Caro, E. M. Schmidt, C.-Y. Sha, and Y.-H. Yang, “1000 songs for emotional analysis of music,” in ACM International Workshop on Crowdsourcing for Multimedia , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. van den Oord, S. Dieleman, and B. Schrauwen, “Transfer learning by supervised pre-training for audio-based music classification,” in ISMIR , 2014
2014
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in European Conference on Computer Vision , 2014
2014
Earlier work this paper cites.
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, and D. P. Ellis, “mir_eval: A transparent implementation of common mir metrics,” in ISMIR , 2014
2014
Earlier work this paper cites.
F. Weninger, F. Eyben, and B. Schuller, “On-line continuous-time music mood regression with deep recurrent neural networks,” in ICASSP , 2014
2014
Earlier work this paper cites.
C. Kereliuk, B. L. Sturm, and J. Larsen, “Deep learning and music adversaries,” IEEE Transactions on Multimedia , 2015
2015
Earlier work this paper cites.
P. Knees, Á. Faraldo Pérez, H. Boyer, R. Vogl, S. Böck, F. Hörschläger, M. Le Goff et al. , “Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections,” in ISMIR , 2015
2015
Earlier work this paper cites.
Pioneer, “rekordbox v3.2.2,” 2015. [Online]. Available: http://www.cp.jku.at/datasets/giantsteps/
2015
Earlier work this paper cites.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Python in Science Conference , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. Choi, G. Fazekas, M. Sandler, and K. Cho, “Transfer learning for music classification and regression tasks,” in ISMIR , 2017
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
F. Korzeniowski and G. Widmer, “End-to-end musical key estimation using a convolutional neural network,” in European Signal Processing Conference , 2017
2017
Cited alongside, same era.
F. Medhat, D. Chesmore, and J. Robinson, “Masked conditional neural networks for audio classification,” in International Conference on Artificial Neural Networks , 2017
J. Hewitt and C. D. Manning, “A structural probe for finding syntax in word representations,” in North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019
2019
Later among the works it cites.
J. Jiang, G. G. Xia, and D. B. Carlton, “MIREX 2019 submission: Crowd annotation for audio key estimation,” MIREX , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music transformer,” in ICLR , 2019
2019
Later among the works it cites.
C. Donahue, H. H. Mao, Y. E. Li, G. W. Cottrell, and J. McAuley, “LakhNES: Improving multi-instrumental music generation with cross-domain pre-training,” in ISMIR , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” in ICLR , 2017
2017
Cited alongside, same era.
I. Simon and S. Oore, “Performance RNN: Generating music with expressive timing and dynamics,” 2017. [Online]. Available: https://magenta.tensorflow.org/performance-rnn
2017
Cited alongside, same era.
J. Lee, J. Park, K. L. Kim, and J. Nam, “SampleCNN: End-to-end deep convolutional neural networks using very small filters for music classification,” Applied Sciences , 2018
2018
Cited alongside, same era.
B. McFee, J. W. Kim, M. Cartwright, J. Salamon, R. M. Bittner, and J. P. Bello, “Open-source practices for music signal processing research: Recommendations for transparent, sustainable, and reproducible audio research,” IEEE Signal Processing Magazine , 2018
2018
Cited alongside, same era.
S. Dieleman, A. van den Oord, and K. Simonyan, “The challenge of realistic music generation: modelling raw audio at scale,” in NIPS , 2018
2018
Cited alongside, same era.
D. Hupkes, S. Veldhoen, and W. Zuidema, “Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure,” Journal of Artificial Intelligence Research , 2018
2018
Cited alongside, same era.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
H.-T. Hung, C.-Y. Wang, Y.-H. Yang, and H.-M. Wang, “Improving automatic jazz melody generation by transfer learning techniques,” in Asia-Pacific Signal and Information Processing Association Annual Summit and Conference , 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
Q. Huang, A. Jansen, L. Zhang, D. P. Ellis, R. A. Saurous, and J. Anderson, “Large-scale weakly-supervised content embeddings for music recommendation and tagging,” in ICASSP , 2020
2020
Later among the works it cites.
J. Kim, J. Urbano, C. C. Liem, and A. Hanjalic, “One deep music representation to rule them all? A comparative analysis of different representation learning strategies,” Neural Computing and Applications , 2020
2020
Later among the works it cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML , 2020
2020
Later among the works it cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in ICML , 2020
2020
Later among the works it cites.
E. A. Chi, J. Hewitt, and C. D. Manning, “Finding universal grammatical relations in multilingual bert,” in Association for Computational Linguistics , 2020
2020
Later among the works it cites.
A. Rogers, O. Kovaleva, and A. Rumshisky, “A primer in BERTology: What we know about how BERT works,” Transactions of the Association for Computational Linguistics , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
A. Ferraro, X. Favory, K. Drossos, Y. Kim, and D. Bogdanov, “Enriched music representations with multiple cross-modal contrastive learning,” IEEE Signal Processing Letters , 2021
2021
Closest in time.