Fetching the paper…
Reading the bibliography…
We present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT to yield the best of both designs.
Object recognition with gradient-based learning
Yann Lecun, Patrick Haffner, Léon Bottou, and Yoshua Bengio · 1999
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
François Chollet · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J. Corso, Richard Socher, and Caiming Xiong · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Local relation networks for image recognition
Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin · 2019
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
Rethinking transformer-based set prediction for object detection
Zhiqing Sun, Shengcao Cao, Yiming Yang, and Kris Kitani · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
Max-deeplab: End-to-end panoptic segmentation with mask transformers
Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen · 2020
Later among the works it cites.
End-to-end video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli · 2019
Cited alongside, same era.
Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Pre-trained image processing transformer
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao · 2020
Cited alongside, same era.
Da-nas: Data adapted pruning for efficient neural architecture search
Xiyang Dai, Dongdong Chen, Mengchen Liu, Yinpeng Chen, and Lu YUan · 2020
Cited alongside, same era.
Up-detr: Unsupervised pre-training for object detection with transformers
Zhigang Dai, Bolun Cai, Yugeng Lin, and Junying Chen · 2020
Cited alongside, same era.
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han · 2020
Later among the works it cites.
Learning texture transformer network for image super-resolution
Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo · 2020
Later among the works it cites.
Learning joint spatial-temporal transformations for video inpainting
Yanhong Zeng, Jianlong Fu, and Hongyang Chao · 2020
Later among the works it cites.
End-to-end object detection with adaptive clustering transformer
Minghang Zheng, Peng Gao, Xiaogang Wang, Hongsheng Li, and Hao Dong · 2020
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Later among the works it cites.
Do we really need explicit position encodings for vision transformers?
Xiangxiang Chu, Bo Zhang, Zhi Tian, Xiaolin Wei, and Huaxia Xia · 2021
Closest in time.
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang · 2021
Closest in time.
Bottleneck transformers for visual recognition
Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani · 2021
Closest in time.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Closest in time.
Evolving attention with residual convolutions
Yujing Wang, Yaming Yang, Jiangang Bai, Mingliang Zhang, Jing Bai, Jing Yu, Ce Zhang, Gao Huang, and Yunhai Tong · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.