Fetching the paper…
Reading the bibliography…
The transformer architectures with attention mechanisms have obtained success in Nature Language Processing (NLP), and Vision Transformers (ViTs) have recently extended the application domains to various vision tasks.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Earlier work this paper cites.
Fully quantized transformer for improved translation
G. Prato, E. Charlaix, and M. Rezagholizadeh · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu · 2019
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat · 2019
Earlier work this paper cites.
Binarybert: Pushing the limit of bert quantization
H. Bai, W. Zhang, L. Hou, L. Shang, J. Jin, X. Jiang, Q. Liu, M. Lyu, and I. King · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
T. Chen, J. Frankle, S. Chang, S. Liu, Y. Zhang, Z. Wang, and M. Carbin · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Ftrans: energy-efficient acceleration of transformers using fpga
B. Li, S. Pandey, H. Fang, Y. Lyv, J. Li, J. Chen, M. Xie, L. Wan, H. Liu, and C. Ding · 2020
Earlier work this paper cites.
End-to-end human pose and mesh reconstruction with transformers
K. Lin, L. Wang, and Z. Liu · 2020
Earlier work this paper cites.
Reactnet: Towards precise binary neural network with generalized activation functions
Z. Liu, Z. Shen, M. Savvides, and K.-T. Cheng · 2020
Cited alongside, same era.
Q-bert: Hessian based ultra low precision quantization of bert
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer · 2020
Cited alongside, same era.
End-to-end video instance segmentation with transformers
Y. Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, and H. Xia · 2020
Cited alongside, same era.
Ternarybert: Distillation-aware ultra-low bit bert
W. Zhang, L. Hou, Y. Yin, L. Shang, X. Chen, X. Jiang, and Q. Liu · 2020
Cited alongside, same era.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr, et al · 2020
Efficient training of visual transformers with small-size datasets
Y. Liu, E. Sangineto, W. Bi, N. Sebe, B. Lepri, and M. De Nadai · 2021
Later among the works it cites.
Hardware acceleration of fully quantized bert for efficient natural language processing
Z. Liu, G. Li, and J. Cheng · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Post-training quantization for vision transformer
Z. Liu, Y. Wang, K. Han, S. Ma, and W. Gao · 2021
Later among the works it cites.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2021
Cited alongside, same era.
Crossvit: Cross-attention multi-scale vision transformer for image classification
C.-F. Chen, Q. Fan, and R. Panda · 2021
Cited alongside, same era.
Autoformer: Searching transformers for visual recognition
M. Chen, H. Peng, J. Fu, and H. Ling · 2021
Cited alongside, same era.
Chasing sparsity in vision transformers: An end-to-end exploration
T. Chen, Y. Cheng, Z. Gan, L. Yuan, L. Zhang, and Z. Wang · 2021
Cited alongside, same era.
When vision transformers outperform resnets without pretraining or strong data augmentations
X. Chen, C.-J. Hsieh, and B. Gong · 2021
Cited alongside, same era.
Xcit: Cross-covariance image transformers
A. El-Nouby, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, et al · 2021
Cited alongside, same era.
S. Mehta and M. Rastegari · 2021
Later among the works it cites.
Accelerating transformer-based deep learning models on fpgas using column balanced block pruning
H. Peng, S. Huang, T. Geng, A. Li, W. Jiang, H. Liu, S. Wang, and C. Ding · 2021
Later among the works it cites.
Accommodating transformer onto fpga: Coupling the balanced model compression and fpga-implementation optimization
P. Qi, Y. Song, H. Peng, S. Huang, Q. Zhuge, and E. H.-M. Sha · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers, 2021
A. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J’egou · 2021
Later among the works it cites.
Kvt: k-nn attention for boosting vision transformers, 2021
P. Wang, X. Wang, F. Wang, M. Lin, S. Chang, W. Xie, H. Li, and R. Jin · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Rethinking and improving relative position encoding for vision transformer, 2021
K. Wu, H. Peng, M. Chen, J. Fu, and H. Chao · 2021
Later among the works it cites.
Incorporating convolution designs into visual transformers, 2021
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z. Jiang, F. E. Tay, J. Feng, and S. Yan · 2021
Later among the works it cites.
Vision transformer with progressive sampling
X. Yue, S. Sun, Z. Kuang, M. Wei, P. Torr, W. Zhang, and D. Lin · 2021
Later among the works it cites.