Fetching the paper…
Reading the bibliography…
Vision transformer (ViT) is an attention neural network architecture that is shown to be effective for computer vision tasks.
A kernel statistical test of independence
Arthur Gretton, Kenji Fukumizu, Choon Teo, Le Song, Bernhard Schölkopf, and Alex Smola · 2007
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Earlier work this paper cites.
Experiment tracking with weights and biases
Lukas Biewald · 2020
Cited alongside, same era.
Hybrid byol-vit: Efficient approach to deal with small datasets
Safwen Naimi, Rien van Leeuwen, Wided Souidene, and Slim Ben Saoud · 2021
Cited alongside, same era.
Vision transformer for small-size datasets
Seung Hoon Lee, Seunghyun Lee, and Byung Cheol Song · 2021
Cited alongside, same era.
Do vision transformers see like convolutional neural networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Efficient training of visual transformers with small datasets
Yahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, and Marco Nadai · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…