Fetching the paper…
Reading the bibliography…
Recently, neural networks based purely on self-attention, such as the Vision Transformer (ViT), have been shown to outperform deep learning models constructed with convolutional neural networks (CNNs) on various vision tasks, thus extending the success of Transformers, which were originally developed for language processing, to the vision domain.
Learning problem-agnostic speech representations from multiple self-supervised tasks
Pascual, S.; Ravanelli, M.; Serra, J.; Bonafonte, A.; and Bengio, Y. 2019 · 1904
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Schneider, S.; Baevski, A.; Collobert, R.; and Auli, M. 2019 · 1904
Earlier work this paper cites.
Selfie: Self-supervised pretraining for image embedding
Trinh, T. H.; Luong, M.-T.; and Le, Q. V. 2019 · 1906
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
LeCun, Y.; and Bengio, Y. 1995 · 1995
Earlier work this paper cites.
IEMOCAP: Interactive emotional dyadic motion capture database
Busso, C.; Bulut, M.; Lee, C.-C.; Kazemzadeh, A.; Mower, E.; Kim, S.; Chang, J. N.; Lee, S.; and Narayanan, S. S. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Jiang, D.; Li, W.; Cao, M.; Zhang, R.; Zou, W.; Han, K.; and Li, X. 2020 · 2010
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2020 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V.; Chen, G.; Povey, D.; and Khudanpur, S. 2015 · 2015
Earlier work this paper cites.
ESC: Dataset for environmental sound classification
Piczak, K. J. 2015 · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M.; and Favaro, P. 2016 · 2016
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Cited alongside, same era.
Learning from between-class examples for deep sound recognition
Tokozume, Y.; Ushiku, Y.; and Harada, T. 2018 · 2018
Cited alongside, same era.
Speech commands: A dataset for limited-vocabulary speech recognition
Warden, P. 2018 · 2018
Cited alongside, same era.
Sit: Self-supervised vision transformer
Atito, S.; Awais, M.; and Kittler, J. 2021 · 2021
Closest in time.
BEiT: BERT Pre-Training of Image Transformers
Bao, H.; Dong, L.; and Wei, F. 2021 · 2021
Closest in time.
Keyword Transformer: A Self-Attention Model for Keyword Spotting
Berg, A.; O’Connor, M.; and Cruz, M. T. 2021 · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Closest in time.
An empirical study of training self-supervised vision transformers
Chen, X.; Xie, S.; and He, K. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An unsupervised autoregressive model for speech representation learning
Chung, Y.-A.; Hsu, W.-N.; Tang, H.; and Glass, J. 2019 · 2019
Cited alongside, same era.
Bag of tricks for image classification with convolutional neural networks
He, T.; Zhang, Z.; Zhang, H.; Zhang, Z.; Xie, J.; and Li, M. 2019 · 2019
Cited alongside, same era.
SpecAugment: A simple data augmentation method for automatic speech recognition
Park, D. S.; Chan, W.; Zhang, Y.; Chiu, C.-C.; Zoph, B.; Cubuk, E. D.; and Le, Q. V. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020 · 2020
Cited alongside, same era.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Liu, A. T.; Yang, S.-w.; Chi, P.-H.; Hsu, P.-c.; and Lee, H.-y. 2020 · 2020
Cited alongside, same era.
Voxceleb: Large-scale speaker verification in the wild
Nagrani, A.; Chung, J. S.; Xie, W.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
Multi-task self-supervised learning for robust speech recognition
Ravanelli, M.; Zhong, J.; Pascual, S.; Swietojanski, P.; Monteiro, J.; Trmal, J.; and Bengio, Y. 2020 · 2020
Cited alongside, same era.
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Closest in time.
AST: Audio Spectrogram Transformer
Gong, Y.; Chung, Y.-A.; and Glass, J. 2021 · 2021
Closest in time.
HuBERT: How much can a bad teacher benefit ASR pre-training?
Hsu, W.-N.; Tsai, Y.-H. H.; Bolte, B.; Salakhutdinov, R.; and Mohamed, A. 2021 · 2021
Closest in time.
BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
Niizumi, D.; Takeuchi, D.; Ohishi, Y.; Harada, N.; and Kashino, K. 2021 · 2021
Closest in time.
Contrastive learning of general-purpose audio representations
Saeed, A.; Grangier, D.; and Zeghidour, N. 2021 · 2021
Closest in time.
SUPERB: Speech processing Universal PERformance Benchmark
Yang, S.-w.; Chi, P.-H.; Chuang, Y.-S.; Lai, C.-I. J.; Lakhotia, K.; Lin, Y. Y.; Liu, A. T.; Shi, J.; Chang, X.; Lin, G.-T.; et al. 2021 · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.; Tay, F. E.; Feng, J.; and Yan, S. 2021 · 2021
Closest in time.