Fetching the paper…
Reading the bibliography…
Transformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships.
L. K. Saul and M. G. Rahim, “Maximum likelihood and minimum classification error factor analysis for automatic speech recognition,” IEEE Transactions on Speech and Audio Processing , vol. 8, no. 2, pp. 115–125, 2000
2000
Earlier work this paper cites.
S. J. D. Prince, J. H. Elder, J. Warrell, and F. M. Felisberti, “Tied factor analysis for face recognition across large pose differences,” IEEE Transactions on Pattern Analysis & Machine Intelligence , vol. 30, no. 6, pp. 970–984, 2008
2008
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proceedings of the 25th ICML , 2008, pp. 1096–1103
2008
Earlier work this paper cites.
L. Rabiner and R. Schafer, Theory and applications of digital speech processing . Prentice Hall Press, 2010
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” NeurIPS , vol. 25, pp. 1097–1105, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. J. Piczak, “Esc: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM International Conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in ICASSP . IEEE, 2017, pp. 131–135
2017
Earlier work this paper cites.
Y. Xu, Q. Huang, W. Wang, P. Foster, S. Sigtia, P. Jackson, and M. Plumbley, “Unsupervised feature learning based on deep models for environmental audio tagging,” IEEE/ACM Transactions on Audio Speech and Language Processing , vol. 25, pp. 1230–1241, 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE ICASSP . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 3733–3742
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Kendall, Y. Gal, and R. Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 7482–7491
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio, “Learning deep representations by mutual information estimation and maximization,” in ICLR , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , vol. 1, 2019, p. 2
2019
Cited alongside, same era.
H. Al-Tahan and Y. Mohsenzadeh, “CLAR: Contrastive learning of auditory representations,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 2530–2538
2021
Later among the works it cites.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for audio: Self-supervised learning for general-purpose audio representation,” in 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2021, pp. 1–8
2021
Later among the works it cites.
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny, “Barlow Twins: Self-supervised learning via redundancy reduction,” in ICML . PMLR, 2021, pp. 12 310–12 320
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML . PMLR, 2020, pp. 1597–1607
2020
Cited alongside, same era.
M. Patacchiola and A. Storkey, “Self-supervised relational reasoning for representation learning,” in NeurIPS , vol. 33, 2020, pp. 4003–4014
2020
Cited alongside, same era.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” NeurIPS , vol. 33, pp. 21 271–21 284, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
M. BRiviere, A. Joulin, P.-E. Mazare, and D. Emmanuel, “Unsupervised pretraining transfers well across languages,” ICASSP , 2020
2020
Cited alongside, same era.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 9650–9660
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International conference on machine learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
2021
Later among the works it cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in ICASSP . IEEE, 2015, pp. 5206–5210
2021
Later among the works it cites.
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “SSAST: Self-supervised audio spectrogram transformer,” in Proceedings of the AAAI , vol. 36, no. 10, 2022, pp. 10 699–10 709
2022
Closest in time.
2022
Closest in time.
S. Atito, S. M. Anwar, M. Awais, and J. Kittler, “SB-SSL: Slice-based self-supervised transformers for knee abnormality classification from MRI,” in Workshop on Medical Image Learning with Limited and Noisy Data, MICCAI . Springer, 2022, pp. 86–95
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked autoencoders that listen,” NeurIPS , vol. 35, pp. 28 708–28 720, 2022
2022
Closest in time.
S. Atito, M. Awais, and J. Kittler, “GMML is all you need,” arXiv preprint arXiv:2205.14986 , 2022
2022
Closest in time.
H. Bao, L. Dong, S. Piao, and F. Wei, “BEit: BERT pre-training of image transformers,” in ICLR , 2022
2022
Closest in time.
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “SimMIM: A simple framework for masked image modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 9653–9663
2022
Closest in time.
D. Chong, H. Wang, P. Zhou, and Q. Zeng, “Masked spectrogram prediction for self-supervised audio pre-training,” in ICASSP . IEEE, 2023, pp. 1–5
2023
Closest in time.