Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has emerged as a popular approach for learning audio representations.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proceedings of the 22nd ACM international conference on Multimedia , 2014, pp. 1041–1044
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
——, “Sgdr - stochastic gradient descent with warm restarts,” arXiv:1608.03983 , 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv:1711.05101 [cs] , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with wavenet autoencoders,” in International Conference on Machine Learning . PMLR, 2017, pp. 1068–1077
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. JiaKai, “Mean teacher convolution system for dcase 2018 task 4,” DCASE2018 Challenge, Tech. Rep., June 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Turpault, R. Serizel, A. Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Acoustic Scenes and Events 2019 Workshop (DCASE2019) , 2019, p. 253
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Tagliasacchi, B. Gfeller, F. de Chaumont Quitry, and D. Roblek, “Pre-Training Audio Representations With Self-Supervision,” IEEE Signal Processing Letters , vol. 27, pp. 600–604, 2020
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
X. Chen and K. He, “Exploring Simple Siamese Representation Learning,” arXiv:2011.10566 [cs] , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2022
Later among the works it cites.
A. Baade, P. Peng, and D. Harwath, “MAE-AST: Masked Autoencoding Audio Spectrogram Transformer,” Mar. 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei, “BEATs: Audio Pre-Training with Acoustic Tokenizers,” Dec. 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee, “Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6419–6423, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Ç. Bilen, G. Ferroni, F. Tuveri, J. Azcarreta, and S. Krstulović, “A framework for the robust evaluation of sound event detection,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 61–65
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra, “Unsupervised Contrastive Learning of Sound Event Representations,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 371–375
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
X. LI and X. Li, “ATST: Audio Representation Learning with Teacher-Student Transformer,” in Proc. Interspeech 2022 , 2022, pp. 4172–4176
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language,” Oct. 2022
2022
Later among the works it cites.
S. Atito, M. Awais, W. Wang, M. D. Plumbley, and J. Kittler, “ASiT: Audio Spectrogram vIsion Transformer for General Audio Representation,” Nov. 2022
2022
Later among the works it cites.
Y. Gong, S. Khurana, A. Rouditchenko, and J. Glass, “CMKD: CNN/Transformer-Based Cross-Model Knowledge Distillation for Audio Classification,” Mar. 2022
2022
Later among the works it cites.
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “FSD50K: An Open Dataset of Human-Labeled Sound Events,” Apr. 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Wang, P. Luc, Y. Wu, A. Recasens, L. Smaira, A. Brock, A. Jaegle, J.-B. Alayrac, S. Dieleman, J. Carreira, and A. v. d. Oord, “Towards Learning Universal Audio Representations,” Jun. 2022
2022
Later among the works it cites.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Singapore, Singapore: IEEE, May 2022, pp. 646–650
2022
Later among the works it cites.
K. Koutini, J. Schlüter, H. Eghbal-zadeh, and G. Widmer, “Efficient Training of Audio Transformers with Patchout,” in Interspeech 2022 . ISCA, Sep. 2022, pp. 2753–2757
2022
Later among the works it cites.
J. Turian, J. Shier, H. R. Khan, B. Raj, B. W. Schuller, C. J. Steinmetz, C. Malloy, G. Tzanetakis, G. Velarde, K. McNally, M. Henry, N. Pinto, C. Noufi, C. Clough, D. Herremans, E. Fonseca, J. Engel, J. Salamon, P. Esling, P. Manocha, S. Watanabe, Z. Jin, and Y. Bisk, “HEAR: Holistic Evaluation of Audio Representations,” May 2022
2022
Later among the works it cites.
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked Autoencoders that Listen,” Jan. 2023
2023
Closest in time.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 137–151, 2023
2023
Closest in time.
——, “Masked modeling duo: Learning representations by encouraging both networks to model the input,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.
D. Chong, H. Wang, P. Zhou, and Q. Zeng, “Masked spectrogram prediction for self-supervised audio pre-training,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.