Fetching the paper…
Reading the bibliography…
The Position Embedding (PE) is critical for Vision Transformers (VTs) due to the permutation-invariance of self-attention operation.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Earlier work this paper cites.
Attention augmented convolutional networks
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
Segmentation transformer: Object-contextual representations for semantic segmentation
Yuhui Yuan, Xiaokang Chen, Xilin Chen, and Jingdong Wang · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Earlier work this paper cites.
On position embeddings in bert
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen · 2020
Earlier work this paper cites.
On position embeddings in bert
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen · 2020
Earlier work this paper cites.
Yu-An Wang and Yun-Nung Chen · 2020
Cited alongside, same era.
Adaptive image transformer for one-shot object detection
Ding-Jie Chen, He-Yen Hsieh, and Tyng-Luh Liu · 2021
Cited alongside, same era.
Conditional positional encodings for vision transformers
Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Cited alongside, same era.
Convit: Improving vision transformers with soft convolutional inductive biases
Stéphane d’Ascoli, Hugo Touvron, Matthew L Leavitt, Ari S Morcos, Giulio Biroli, and Levent Sagun · 2021
Cited alongside, same era.
Escaping the big data paradigm with compact transformers
Ali Hassani, Steven Walton, Nikhil Shah, Abulikemu Abuduweili, Jiachen Li, and Humphrey Shi · 2021
Cited alongside, same era.
Incorporating convolution designs into visual transformers
Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Recurrent glimpse-based decoder for detection with transformer
Zhe Chen, Jing Zhang, and Dacheng Tao · 2022
Closest in time.
Position information in transformers: An overview
Philipp Dufter, Martin Schmitt, and Hinrich Schütze · 2022
Closest in time.
Embracing single stride 3d object detector with sparse transformer
Lue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang, Hang Zhao, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang · 2022
Closest in time.
Voxel set transformer: A set-to-set approach to 3d object detection from point clouds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rethinking spatial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh · 2021
Cited alongside, same era.
Learning position and target consistency for memory-based video object segmentation
Li Hu, Peng Zhang, Bang Zhang, Pan Pan, Yinghui Xu, and Rong Jin · 2021
Cited alongside, same era.
Hcrf-flow: Scene flow from point clouds with continuous high-order crfs and position-aware flow embedding
Ruibo Li, Guosheng Lin, Tong He, Fayao Liu, and Chunhua Shen · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
2lspe: 2d learnable sinusoidal positional encoding using transformer for scene text recognition
Zobeir Raisi, Mohamed A Naiel, Georges Younes, Steven Wardell, and John Zelek · 2021
Cited alongside, same era.
Segmenter: Transformer for semantic segmentation
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Cited alongside, same era.
Chenhang He, Ruihuang Li, Shuai Li, and Lei Zhang · 2022
Closest in time.
Destr: Object detection with split transformer
Liqiang He and Sinisa Todorovic · 2022
Closest in time.
Expectation-maximization contrastive learning for compact video-and-language representations
Peng Jin, JinFa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David A. Clifton, and Jie Chen · 2022
Closest in time.
Toward 3d spatial reasoning for human-like text-based visual question answering
Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen · 2022
Closest in time.
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo · 2022
Closest in time.
Towards robust vision transformer
Xiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li, Ranjie Duan, Shaokai Ye, Yuan He, and Hui Xue · 2022
Closest in time.
Bridged transformer for vision and point cloud 3d object detection
Yikai Wang, TengQi Ye, Lele Cao, Wenbing Huang, Fuchun Sun, Fengxiang He, and Dacheng Tao · 2022
Closest in time.
Uformer: A general u-shaped transformer for image restoration
Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, and Houqiang Li · 2022
Closest in time.
Multi-class token transformer for weakly supervised semantic segmentation
Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu · 2022
Closest in time.
Volo: Vision outlooker for visual recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan · 2022
Closest in time.
Nested hierarchical transformer: Towards accurate, data-efficient and interpretable visual understanding
Zizhao Zhang, Han Zhang, Long Zhao, Ting Chen, Sercan Ö Arik, and Tomas Pfister · 2022
Closest in time.