Fetching the paper…
Reading the bibliography…
Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features.
musicnn: Pre-trained convolutional neural networks for music audio tagging
Pons, J. and Serra, X · 1909
Earlier work this paper cites.
Musical genre classification of audio signals
Tzanetakis, G. and Cook, P · 2002
Earlier work this paper cites.
Canonical correlation analysis: An overview with application to learning methods
Hardoon, D. R., Szedmak, S., and Shawe-Taylor, J · 2004
Earlier work this paper cites.
The million song dataset
Bertin-Mahieux, T., Ellis, D. P., Whitman, B., and Lamere, P · 2011
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Freesound technical demo
Font, F., Roma, G., and Serra, X · 2013
Earlier work this paper cites.
Cross-modal retrieval with correspondence autoencoder
Feng, F., Wang, X., and Li, R · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Salamon, J., Jacoby, C., and Bello, J. P · 2014
Earlier work this paper cites.
Learning grounded meaning representations with autoencoders
Silberer, C. and Lapata, M · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Transfer learning by supervised pre-training for audio-based music classification
Van Den Oord, A., Dieleman, S., and Schrauwen, B · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Deep learning and music adversaries
Kereliuk, C., Sturm, B. L., and Larsen, J · 2015
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., and Nieto, O · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S · 2016
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Aytar, Y., Vondrick, C., and Torralba, A · 2016
Cited alongside, same era.
A guide to convolution arithmetic for deep learning, 2016
Dumoulin, V. and Visin, F · 2016
Cited alongside, same era.
The extended ballroom dataset
Marchand, U. and Peeters, G · 2016
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Deep convolutional neural networks and data augmentation for environmental sound classification
Salamon, J. and Bello, J. P · 2017
Later among the works it cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Later among the works it cites.
Mad twinnet: Masker-denoiser architecture with twin networks for monaural sound source separation
Drossos, K., Mimilakis, S. I., Serdyuk, D., Schuller, G., Virtanen, T., and Bengio, Y · 2018
Later among the works it cites.
Facilitating the manual annotation of sounds when using large taxonomies
Favory, X., Fonseca, E., Font, F., and Serra, X · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sohn, K · 2016
Cited alongside, same era.
Sequence to sequence autoencoders for unsupervised representation learning from audio
Amiriparian, S., Freitag, M., Cummins, N., and Schuller, B · 2017
Cited alongside, same era.
Look, listen and learn
Arandjelovic, R. and Zisserman, A · 2017
Cited alongside, same era.
Transfer learning for music classification and regression tasks
Choi, K., Fazekas, G., Sandler, M., and Cho, K · 2017
Cited alongside, same era.
Neural audio synthesis of musical notes with wavenet autoencoders
Engel, J., Resnick, C., Roberts, A., Dieleman, S., Norouzi, M., Eck, D., and Simonyan, K · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Cited alongside, same era.
Cnn architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., et al · 2017
Cited alongside, same era.
Lee, J., Park, J., Kim, K. L., and Nam, J · 2018
Later among the works it cites.
Monaural singing voice separation with skip-filtering connections and recurrent inference of time-frequency mask
Mimilakis, S. I., Drossos, K., Santos, J. F., Schuller, G., Virtanen, T., and Bengio, Y · 2018
Later among the works it cites.
Look, listen, and learn more: Design choices for deep audio embeddings
Cramer, J., Wu, H.-H., Salamon, J., and Bello, J. P · 2019
Later among the works it cites.
Randomly weighted cnns for (music) audio classification
Pons, J. and Serra, X · 2019
Later among the works it cites.
Data augmentation for instrument classification robust to audio effects
Ramires, A. and Serra, X · 2019
Later among the works it cites.
Generalized zero-and few-shot learning via aligned variational autoencoders
Schonfeld, E., Ebrahimi, S., Sinha, S., Darrell, T., and Akata, Z · 2019
Later among the works it cites.
Tensorflow audio models in essentia
Alonso-Jiménez, P., Bogdanov, D., Pons, J., and Serra, X · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Closest in time.