Fetching the paper…
Reading the bibliography…
While deep learning has enabled great advances in many areas of music, labeled music datasets remain especially hard, expensive, and time-consuming to create.
S. Stevens, J. Volkmann, and E. B. Newman, “A Scale for the Measurement of the Psychological Magnitude Pitch,” Journal of the Acoustical Society of America , vol. 8, pp. 185–190, 1937
1937
Earlier work this paper cites.
1961
Earlier work this paper cites.
M. R. Schroeder, “Natural Sounding Artificial Reverberation,” Journal of the Audio Engineering Society , vol. 10, no. 3, pp. 219–223, July 1962
1962
Earlier work this paper cites.
2002
Earlier work this paper cites.
G. Tzanetakis and P. Cook, “Musical Genre Classification of Audio Signals,” IEEE Transactions on speech and audio processing , vol. 10, no. 5, pp. 293–302, 2002
2002
Earlier work this paper cites.
U. Zölzer, X. Amatriain, D. Arfib, J. Bonada, G. De Poli, P. Dutilleux, G. Evangelista, F. Keiler, A. Loscos, D. Rocchesso et al. , DAFX-Digital Audio Effects . John Wiley & Sons, 2002
2002
Earlier work this paper cites.
J. Davis and M. Goadrich, “The Relationship between Precision-Recall and ROC Curves,” in Proceedings of the 23rd International Conference on Machine Learning , ser. ICML ’06. New York, NY, USA: Association for Computing Machinery, 2006, p. 233–240. [Online]. Available: https://doi.org/10.1145/1143844.1143874
2006
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing Data using t-SNE,” Journal of machine learning research , vol. 9, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of Algorithms Using Games: The Case of Music Tagging,” in Proceedings of the 10th International Society for Music Information Retrieval Conference , 2009
2009
Earlier work this paper cites.
P. Hamel, S. Lemieux, Y. Bengio, and D. Eck, “Temporal Pooling and Multiscale Learning for Automatic Annotation and Ranking of Music Audio,” in Proceedings of the 12th International Society for Music Information Retrieval Conference, ISMIR , 2011, pp. 729–734
2011
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The Million Song Dataset,” in Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR 2011) , 2011
2011
Earlier work this paper cites.
J. A. Burgoyne, J. Wild, and I. Fujinaga, “An Expert Ground Truth Set for Audio Chord Recognition and Music Analysis,” in Proceedings of the 12th International Society for Music Information Retrieval Conference, ISMIR , 2011
2011
Earlier work this paper cites.
A. van den Oord, S. Dieleman, and B. Schrauwen, “Deep content-based music recommendation,” in Advances in Neural Information Processing Systems , C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, Eds., vol. 26. Curran Associates, Inc., 2013, pp. 2643–2651. [Online]. Available: https://proceedings.neurips.cc/paper/2013/file/b3ba8f1bee1238a2f37603d90b58898d-Paper.pdf
2013
Earlier work this paper cites.
S. Dieleman and B. Schrauwen, “Multiscale Approaches to Music Audio Feature Learning,” in Proceedings of the 14th International Society for Music Information Retrieval conference , 2013, pp. 116–121
2013
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation Learning: A Review and New Perspectives,” IEEE transactions on pattern analysis and machine intelligence , vol. 35, no. 8, pp. 1798–1828, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
D. Bogdanov, N. Wack, E. Gómez, S. Gulati, P. Herrera, O. Mayor, G. Roma, J. Salamon, J. R. Zapata, and X. Serra, “ESSENTIA: An Audio Analysis Library for Music Information Retrieval,” in International Society for Music Information Retrieval Conference (ISMIR’13) , Curitiba, Brazil, 04/11/2013 2013, pp. 493–498. [Online]. Available: http://hdl.handle.net/10230/32252
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative Adversarial Nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Cited alongside, same era.
S. Dieleman and B. Schrauwen, “End-to-End Learning for Music Audio,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 6964–6968
2014
Cited alongside, same era.
A. Dosovitskiy, P. Fischer, J. T. Springenberg, M. Riedmiller, and T. Brox, “Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 38, no. 9, pp. 1734–1747, 2015
2015
Cited alongside, same era.
C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised Visual Representation Learning by Context Prediction,” in 2015 IEEE International Conference on Computer Vision (ICCV) . IEEE, 2015, pp. 1422–1430. [Online]. Available: http://ieeexplore.ieee.org/document/7410524/
J. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, “Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 3852–3856
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
S. Böck, F. Krebs, and G. Widmer, “Joint Beat and Downbeat Tracking with Recurrent Neural Networks,” in Proceedings of the 17th International Society for Music Information Retrieval Conference, ISMIR , 2016
2016
Cited alongside, same era.
R. Zhang, P. Isola, and A. A. Efros, “Colorful Image Colorization,” in European conference on computer vision . Springer, 2016, pp. 649–666
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with wavenet autoencoders,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70 , ser. ICML’17. JMLR.org, 2017, p. 1068–1077
2017
Cited alongside, same era.
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning Problem-Agnostic Speech Representations from Multiple Self-Supervised Tasks,” in Proc. Interspeech 2019 , 2019, pp. 161–165. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-2605
2019
Later among the works it cites.
M. Tagliasacchi, B. Gfeller, F. d. C. Quitry, and D. Roblek, “Pre-Training Audio Representations With Self-Supervision,” IEEE Signal Processing Letters , vol. 27, pp. 600–604, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
B. Gfeller, C. Frank, D. Roblek, M. Sharifi, M. Tagliasacchi, and M. Velimirović, “Pitch Estimation Via Self-Supervision,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 3527–3531
2020
Later among the works it cites.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-Task Self-Supervised Learning for Robust Speech Recognition,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6989–6993, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
H. Al-Tahan and Y. Mohsenzadeh, “CLAR: Contrastive Learning of Auditory Representations,” in Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , ser. Proceedings of Machine Learning Research, A. Banerjee and K. Fukumizu, Eds., vol. 130. PMLR, 13–15 Apr 2021, pp. 2530–2538. [Online]. Available: http://proceedings.mlr.press/v130/al-tahan21a.html
2021
Closest in time.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive Learning of General-Purpose Audio Representations,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 3875–3879
2021
Closest in time.
M. Won, K. Choi, and X. Serra, “Semi-supervised Music Tagging Transformer,” In Proc. of International Society for Music Information Retrieval Conference (ISMIR) , 2021
2021
Closest in time.
J. Spijkervet, “Spijkervet/torchaudio-augmentations,” 2021. [Online]. Available: https://doi.org/10.5281/zenodo.5042440
2021
Closest in time.