Fetching the paper…
Reading the bibliography…
The Softmax attention mechanism in Transformer models is notoriously computationally expensive, particularly due to its quadratic complexity, posing significant challenges in vision applications.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, et al · 2009
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, and Zhuang Liu · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, and Kaiming He andPiotr Dollár · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, et al · 2017
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Zhaowei Cai and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, et al · 2018
Earlier work this paper cites.
MMDetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, et al · 2019
Earlier work this paper cites.
Panoptic feature pyramid networks
Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollár · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?, 2019
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Earlier work this paper cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, et al · 2019
Earlier work this paper cites.
Mmsegmentation, an open source semantic segmentation toolbox, 2020
MMSegmentation Contributors · 2020
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, et al · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franc¸ois Fleuret · 2020
Earlier work this paper cites.
Coatnet: Marrying convolution and attention for all data sizes
Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al · 2021
Earlier work this paper cites.
Xcit: Cross-covariance image transformers
Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Earlier work this paper cites.
Efficient attention: Attention with linear complexities
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, et al · 2021
Cited alongside, same era.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Cited alongside, same era.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Cited alongside, same era.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Cited alongside, same era.
Focal self-attention for local-global interactions in vision transformers
Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao · 2021
Cited alongside, same era.
Maxvit: Multi-axis vision transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li · 2022
Later among the works it cites.
Vision transformer with deformable attention
Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang · 2022
Later among the works it cites.
Focal modulation networks
Jianwei Yang, Chunyuan Li, Xiyang Dai, and Jianfeng Gao · 2022
Later among the works it cites.
Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, and Song Han · 2023
Later among the works it cites.
Conditional positional encodings for vision transformers
Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, and Chunhua Shen · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-scale vision longformer: A new vision transformer for high-resolution image encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Cited alongside, same era.
Hydra attention: Efficient attention with many heads, 2022
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, and Judy Hoffman · 2022
Cited alongside, same era.
RegionViT: Regional-to-Local Attention for Vision Transformers
Chun-Fu (Richard) Chen, Rameswar Panda, and Quanfu Fan · 2022
Cited alongside, same era.
Davit: Dual attention vision transformers
Mingyu Ding, Bin Xiao, Noel Codella, et al · 2022
Cited alongside, same era.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Xiaoyi Dong, Jianmin Bao, Dongdong Chen, et al · 2022
Cited alongside, same era.
Doubly-fused vit: Fuse information from vision transformer doubly with local representation
Li Gao, Dong Nie, Bo Li, and Xiaofeng Ren · 2022
Cited alongside, same era.
Conv2former: A simple transformer-style convnet for visual recognition
Qibin Hou, Cheng-Ze Lu, Ming-Ming Cheng, and Jiashi Feng · 2022
Cited alongside, same era.
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Sucheng ren, xingyi yang, songhua liu, xinchao wang
SG-Former: Self guided Transformer with Evolving Token Reallocation · 2023
Later among the works it cites.
Flatten transformer: Vision transformer using focused linear attention
Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, and Gao Huang · 2023
Later among the works it cites.
Neighborhood attention transformer
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi · 2023
Later among the works it cites.
Global context vision transformers
Ali Hatamizadeh, Hongxu Yin, Greg Heinrich, Jan Kautz, and Pavlo Molchanov · 2023
Later among the works it cites.
Vision transformer with super token sampling
Huaibo Huang, Xiaoqiang Zhou, Jie Cao, Ran He, and Tieniu Tan · 2023
Later among the works it cites.
Scale-aware modulation meet transformer
Weifeng Lin, Ziheng Wu, Jiayu Chen, Jun Huang, and Lianwen Jin · 2023
Later among the works it cites.
Internimage: Exploring large-scale vision foundation models with deformable convolutions
Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al · 2023
Later among the works it cites.
Moat: Alternating mobile convolution and attention brings strong vision models
Chenglin Yang, Siyuan Qiao, Qihang Yu, et al · 2023
Later among the works it cites.
Castling-vit: Compressing self-attention via switching towards linear-angular attention during vision transformer inference
Haoran You, Yunyang Xiong, Xiaoliang Dai, Bichen Wu, Peizhao Zhang, Haoqi Fan, Peter Vajda, and Yingyan Lin · 2023
Later among the works it cites.
Biformer: Vision transformer with bi-level routing attention
Lei Zhu, Xinjiang Wang, Zhanghan Ke, Wayne Zhang, and Rynson Lau · 2023
Later among the works it cites.
Demystify mamba in vision: A linear attention perspective
Dongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han, Yifan Pu, Chunjiang Ge, Jun Song, Shiji Song, Bo Zheng, and Gao Huang · 2024
Closest in time.
Learning correlation structures for vision transformers
Manjin Kim, Paul Hongsuck Seo, Cordelia Schmid, and Minsu Cho · 2024
Closest in time.
Moganet: Multi-order gated aggregation network
Siyuan Li, Zedong Wang, Zicheng Liu, Cheng Tan, Haitao Lin, Di Wu, Zhiyuan Chen, Jiangbin Zheng, and Stan Z. Li · 2024
Closest in time.