Query-key normalization for transformers
Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen · 2020
Later among the works it cites.
A proof that rectified deep neural networks overcome the curse of dimensionality in the numerical approximation of semilinear heat equations
Martin Hutzenthaler, Arnulf Jentzen, Thomas Kruse, and Tuan Anh Nguyen · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Finite neuron method and convergence analysis
Jinchao Xu · 2020
Later among the works it cites.
BEiT: BERT pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei · 2021
Later among the works it cites.
Choose a transformer: Fourier or galerkin
Shuhao Cao · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Original
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Later among the works it cites.
An empirical study of training self-supervised vision transformers
Original
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Later among the works it cites.
Deepgreen: deep learning of green’s functions for nonlinear boundary value problems
Craig R Gin, Daniel E Shea, Steven L Brunton, and J Nathan Kutz · 2021
Later among the works it cites.
Transformer in transformer
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
Original
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Original
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Multigraph transformer for free-hand sketch recognition
Peng Xu, Chaitanya K Joshi, and Xavier Bresson · 2021
Later among the works it cites.
Vision transformers are robust learners
Sayak Paul and Pin-Yu Chen · 2022
Closest in time.