Fetching the paper…
Reading the bibliography…
Vision transformer (ViT) and its variants have achieved remarkable successes in various visual tasks.
Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex
David H Hubel and Torsten N Wiesel · 1962
Earlier work this paper cites.
A visual vocabulary for flower classification
Maria-Elena Nilsback and Andrew Zisserman · 2006
Earlier work this paper cites.
Graph theoretical analysis of magnetoencephalographic functional connectivity in alzheimer’s disease
CJ Stam, W De Haan, ABFJ Daffertshofer, BF Jones, I Manshanden, Anne-Marie van Cappellen van Walsum, Teresa Montez, JPA Verbunt, JC De Munck, BW Van Dijk, et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
The brain as an efficient and robust adaptive learner
Sophie Denève, Alireza Alemi, and Ralph Bourdoukan · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Individual differences in learning social and nonsocial network structures
Steven H Tompson, Ari E Kahn, Emily B Falk, Jean M Vettel, and Danielle S Bassett · 2019
Earlier work this paper cites.
Exploring randomly wired neural networks for image recognition
Saining Xie, Alexander Kirillov, Ross Girshick, and Kaiming He · 2019
Cited alongside, same era.
SELFIE: Refurbishing unclean samples for robust deep learning
Hwanjun Song, Minseok Kim, and Jae-Gil Lee · 2019
Cited alongside, same era.
Super-convergence: Very fast training of neural networks using large learning rates
Leslie N Smith and Nicholay Topin · 2019
Cited alongside, same era.
How humans learn and represent networks
Christopher W Lynn and Danielle S Bassett · 2020
Cited alongside, same era.
Graph structure of neural networks
Jiaxuan You, Jure Leskovec, Kaiming He, and Saining Xie · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Resmlp: Feedforward networks for image classification with data-efficient training
Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, et al · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Later among the works it cites.
Shunted self-attention via multi-scale token aggregation
Sucheng Ren, Daquan Zhou, Shengfeng He, Jiashi Feng, and Xinchao Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Cited alongside, same era.
Metaformer is actually what you need for vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2021
Cited alongside, same era.
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2021
Later among the works it cites.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda · 2021
Later among the works it cites.
Imagenet-21k pretraining for the masses
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Uniformer: Unifying convolution and self-attention for visual recognition
Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao · 2022
Closest in time.