Fetching the paper…
Reading the bibliography…
We present Music Tagging Transformer that is trained with a semi-supervised approach.
T. G. Dietterich, R. H. Lathrop, and T. Lozano-Pérez, “Solving the multiple instance problem with axis-parallel rectangles,” Artificial intelligence , vol. 89, no. 1-2, pp. 31–71, 1997
1997
Earlier work this paper cites.
J. Davis and M. Goadrich, “The relationship between precision-recall and roc curves,” in Proc. of international conference on Machine learning (ICML) , 2006
2006
Earlier work this paper cites.
O. Chapelle, B. Scholkopf, and A. Zien, “Semi-supervised learning,” IEEE Transactions on Neural Networks , vol. 20, no. 3, pp. 542–542, 2009
2009
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The million song dataset,” In Proc. of the 12th International Conference on Music Information Retrieval (ISMIR) , 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems (NIPS) , 2012
2012
Earlier work this paper cites.
D.-H. Lee et al. , “Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,” in Workshop on challenges in representation learning, ICML , 2013
2013
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on machine learning (ICML) , 2015
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” Advances in neural information processing systems (NIPS), Deep learning workshop , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
K. Choi, G. Fazekas, and M. Sandler, “Automatic tagging using deep convolutional neural networks,” in In Proc. of the 17th International Society for Music Information Retrieval Conference (ISMIR) , New York, USA, 2016
2016
Earlier work this paper cites.
Y. Kim and A. M. Rush, “Sequence-level knowledge distillation,” Proc. of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. of the IEEE conference on computer vision and pattern recognition (CVPR) , 2016
2016
Earlier work this paper cites.
J. Lee, J. Park, K. L. Kim, and J. Nam, “Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms,” In Proc. of the 14th Sound and music computing (SMC) , 2017
2017
Earlier work this paper cites.
K. Choi, G. Fazekas, M. Sandler, and K. Cho, “Convolutional recurrent neural networks for music classification,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 2392–2396
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Conference on Neural Information Processing Systems (NIPS) , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Pons, O. Nieto, M. Prockup, E. Schmidt, A. Ehmann, and X. Serra, “End-to-end learning for music audio tagging at scale,” in In Proc. of the 19th International Society for Music Information Retrieval Conference (ISMIR) , Paris, France, 2018
2018
Cited alongside, same era.
M. Won, A. Ferraro, D. Bogdanov, and X. Serra, “Evaluation of cnn-based automatic music tagging models,” In Proc. of Sound and Music Computing (SMC) , 2020
2020
Later among the works it cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Later among the works it cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International Conference on Machine Learning (ICML) , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
B. McFee, J. Salamon, and J. P. Bello, “Adaptive pooling operators for weakly labeled sound event detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 11, pp. 2180–2193, 2018
2018
Cited alongside, same era.
K. Choi, G. Fazekas, K. Cho, and M. Sandler, “The effects of noisy labels on deep convolutional neural networks for music tagging,” IEEE Transactions on Emerging Topics in Computational Intelligence , 2018
2018
Cited alongside, same era.
T. Kim, J. Lee, and J. Nam, “Sample-level cnn architectures for music auto-tagging using raw waveforms,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018
2018
Cited alongside, same era.
W. Brendel and M. Bethge, “Approximating cnns with bag-of-local-features models works surprisingly well on imagenet,” International Conference on Learning Representations (ICLR) , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019
2019
Cited alongside, same era.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music transformer,” International Conference on Learning Representations (ICLR) , 2019
2019
Cited alongside, same era.
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le, “Self-training with noisy student improves imagenet classification,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Later among the works it cites.
T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. Hinton, “Big self-supervised models are strong semi-supervised learners,” in Conference on Neural Information Processing Systems (NeurIPS) , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Spijkervet and J. A. Burgoyne, “Contrastive learning of musical representations,” In Proc. of International Society for Music Information Retrieval Conference (ISMIR) , 2021
2021
Closest in time.
2021
Closest in time.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” International Conference on Learning Representations (ICLR) , 2021
2021
Closest in time.
J. Spijkervet, “Spijkervet/torchaudio-augmentations,” 2021. [Online]. Available: https://doi.org/10.5281/zenodo.5042440
2021
Closest in time.
M. Won, S. Oramas, O. Nieto, F. Gouyon, and X. Serra, “Multimodal metric learning for tag-based music retrieval,” In Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.