Fetching the paper…
Reading the bibliography…
Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks.
“Musical genre classification of audio signals,”
G. Tzanetakis and P. Cook, · 2002
Earlier work this paper cites.
“Neural audio synthesis of musical notes with wavenet autoencoders,”
J. Engel et al., · 2010
Earlier work this paper cites.
“Efficient estimation of word representations in vector space,”
T. Mikolov, K. Chen, G. Corrado, and J. Dean, · 2013
Earlier work this paper cites.
“Freesound technical demo,”
F. Font, G. Roma, and X. Serra, · 2013
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
J. Salamon, C. Jacoby, and J. P. Bello, · 2014
Earlier work this paper cites.
“Deep learning and music adversaries,”
C. Kereliuk, B. L Sturm, and J. Larsen, · 2015
Earlier work this paper cites.
“Layer normalization,” 2016
J. L. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Earlier work this paper cites.
“Unsupervised representation learning with deep convolutional generative adversarial networks,”
A. Radford, L. Metz, and S. Chintala, · 2016
Earlier work this paper cites.
“A guide to convolution arithmetic for deep learning,” 2016
V. Dumoulin and F. Visin, · 2016
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
J. F. Gemmeke et al., · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani et al., · 2017
Cited alongside, same era.
“SampleCNN: End-to-end deep convolutional neural networks using very small filters for music classification,”
J. Lee, J. Park, K. L. Kim, and J. Nam, · 2018
Cited alongside, same era.
“Unsupervised learning of semantic audio representations,”
A. Jansen et al., · 2018
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
“Look, listen, and learn more: Design choices for deep audio embeddings,”
J. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, · 2019
Later among the works it cites.
“Self-supervised learning by cross-modal audio-video clustering,”
H. Alwassel, D. Mahajan, L. Torresani, B. Ghanem, and D. Tran, · 2019
Later among the works it cites.
“Semi-supervised triplet loss based learning of ambient audio embeddings,”
N. Turpault, R. Serizel, and E. Vincent, · 2019
Later among the works it cites.
“Learning fragment self-attention embeddings for image-text matching,”
Y. Wu, S. Wang, G. Song, and Q. Huang, · 2019
Later among the works it cites.
“Improving transformer-based speech recognition using unsupervised pre-training,”
D. Jiang et al., · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. van den Oord, Y. Li, and O. Vinyals, · 2018
Cited alongside, same era.
“Mad twinnet: Masker-denoiser architecture with twin networks for monaural sound source separation,”
K. Drossos et al., · 2018
Cited alongside, same era.
“Monaural singing voice separation with skip-filtering connections and recurrent inference of time-frequency mask,”
S. I. Mimilakis et al., · 2018
Cited alongside, same era.
“Musicnn: Pre-trained convolutional neural networks for music audio tagging,”
J. Pons and X. Serra, · 2019
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, · 2020
Closest in time.
“Contrastive representation learning: A framework and review,”
P. H Le-Khac, G. Healy, and A. F Smeaton, · 2020
Closest in time.
“Coala: Co-aligned autoencoders for learning semantically enriched audio representations,”
X. Favory, K. Drossos, T. Virtanen, and X. Serra, · 2020
Closest in time.