Fetching the paper…
Reading the bibliography…
Learning music representations that are general-purpose offers the flexibility to finetune several downstream tasks using smaller datasets.
“Feature learning and deep architectures: New directions for music informatics,”
Eric J Humphrey, Juan P Bello, and Yann LeCun, · 2013
Earlier work this paper cites.
“Transfer learning by supervised pre-training for audio-based music classification,”
Aäron Van Den Oord, Sander Dieleman, and Benjamin Schrauwen, · 2014
Earlier work this paper cites.
“How transferable are features in deep neural networks?,”
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Transfer learning for music classification and regression tasks,”
Keunwoo Choi, György Fazekas, Mark Sandler, and Kyunghyun Cho, · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Learning features of music from scratch,”
J. Thickstun, Z. Harchaoui, and S.M. Kakade, · 2017
Earlier work this paper cites.
“Neural audio synthesis of musical notes with wavenet autoencoders,”
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan, · 2017
Earlier work this paper cites.
“Insights on representational similarity in neural networks with canonical correlation,”
Ari Morcos, Maithra Raghu, and Samy Bengio, · 2018
Cited alongside, same era.
“Crepe: A convolutional representation for pitch estimation,”
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello, · 2018
Cited alongside, same era.
“musicnn: pre-trained convolutional neural networks for music audio tagging,”
Jordi Pons and Xavier Serra, · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Codified audio language modeling learns useful representations for music information retrieval,”
“Contrastive learning of musical representations,”
Janne Spijkervet and John Ashley Burgoyne, · 2021
Later among the works it cites.
“Unsupervised cross-lingual representation learning for speech recognition,”
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli, · 2021
Later among the works it cites.
“Layer-wise analysis of a self-supervised speech representation model,”
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu, · 2021
Later among the works it cites.
“Learning contextual tag embeddings for cross-modal alignment of audio and tags,”
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, and Xavier Serra, · 2021
Later among the works it cites.
“Learning music audio representations via weak language supervision,”
Ilaria Manco, Emmanouil Benetos, Elio Quinton, and György Fazekas, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rodrigo Castellon, Chris Donahue, and Percy Liang, · 2021
Cited alongside, same era.
“Multi-task self-supervised pre-training for music classification,”
Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun, Brian McFee, Juan Pablo Bello, and Chao Wang, · 2021
Cited alongside, same era.
“Musicbert: A self-supervised learning of music representation,”
Hongyuan Zhu, Ye Niu, Di Fu, and Hao Wang, · 2021
Cited alongside, same era.
“Self-supervised learning of audio representations from permutations with differentiable ranking,”
Andrew N Carr, Quentin Berthet, Mathieu Blondel, Olivier Teboul, and Neil Zeghidour, · 2021
Cited alongside, same era.
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, et al., · 2022
Closest in time.
“Exploring the influence of fine-tuning data on wav2vec 2.0 model for blind speech quality prediction,”
Helard Becerra, Alessandro Ragano, and Andrew Hines, · 2022
Closest in time.
“Hear: Holistic evaluation of audio representations,”
Joseph Turian, Jordie Shier, Humair Raj Khan, Bhiksha Raj, Björn W Schuller, Christian J Steinmetz, Colin Malloy, George Tzanetakis, Gissel Velarde, Kirk McNally, et al., · 2022
Closest in time.
“Towards learning universal audio representations,”
Luyu Wang, Pauline Luc, Yan Wu, Adria Recasens, Lucas Smaira, Andrew Brock, Andrew Jaegle, Jean-Baptiste Alayrac, Sander Dieleman, Joao Carreira, et al., · 2022
Closest in time.