Fetching the paper…
Reading the bibliography…
Motivated by biological evolution, this paper explains the rationality of Vision Transformer by analogy with the proven practical evolutionary algorithm (EA) and derives that both have consistent mathematical formulation.
Caltech concurrent computation program, C3P Report 826
Moscato, P., et al.: On evolution, search, optimization, genetic algorithms and martial arts: Towards memetic algorithms · 1989
Earlier work this paper cites.
Cerebral cortex (New York, NY: 1991) (1991)
Felleman, D.J., Van Essen, D.C.: Distributed hierarchical processing in the primate cerebral cortex · 1991
Earlier work this paper cites.
Journal of neurophysiology (1993)
Motter, B.C.: Focal attention produces spatially selective processing in visual cortical areas v1, v2, and v4 in the presence of competing stimuli · 1993
Earlier work this paper cites.
Discrete Applied Mathematics (1994)
Kolen, A., Pesch, E.: Genetic local search in combinatorial optimization · 1994
Earlier work this paper cites.
J. Glob. Optim. pp. 341–359 (1997)
Storn, R., Price, K.: Differential evolution–a simple and efficient heuristic for global optimization over continuous spaces · 1997
Earlier work this paper cites.
University of California, San Diego (1998)
Land, M.W.S.: Evolutionary algorithms with local search for combinatorial optimization · 1998
Earlier work this paper cites.
In: EMO (2003)
Khare, V., Yao, X., Deb, K.: Performance scaling of multi-objective evolutionary algorithms · 2003
Earlier work this paper cites.
Evolutionary computation (2003)
Toffolo, A., Benini, E.: Genetic diversity as an objective in multi-objective evolutionary algorithms · 2003
Earlier work this paper cites.
World Scientific (2004)
Coello, C.A.C., Lamont, G.B.: Applications of multi-objective evolutionary algorithms, vol. 1 · 2004
Earlier work this paper cites.
In: AIAI (2005)
Chen, Z., Kang, L.: Multi-population evolutionary algorithm for solving constrained optimization problems · 2005
Earlier work this paper cites.
In: Recent advances in memetic algorithms, pp. 3–27. Springer (2005)
Hart, W.E., Krasnogor, N., Smith, J.E.: Memetic evolutionary algorithms · 2005
Earlier work this paper cites.
Soft Computing (2005)
Liu, J., Lampinen, J.: A fuzzy adaptive differential evolution algorithm · 2005
Earlier work this paper cites.
TEC (2006)
Brest, J., Greiner, S., Boskovic, B., Mernik, M., Zumer, V.: Self-adapting control parameters in differential evolution: A comparative study on numerical benchmark problems · 2006
Earlier work this paper cites.
In: CEC (2008)
Brest, J., Zamuda, A., Boskovic, B., Maucec, M.S., Zumer, V.: High-dimensional real-parameter optimization using self-adaptive differential evolution algorithm with population size reduction · 2008
Earlier work this paper cites.
In: Advances in metaheuristics for hard optimization. Springer (2008)
García-Martínez, C., Lozano, M.: Local search based on genetic algorithms · 2008
Earlier work this paper cites.
In: CVPR (2009)
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database · 2009
Earlier work this paper cites.
In: CEC (2010)
Brest, J., Zamuda, A., Fister, I., Maučec, M.S.: Large scale global optimization using self-adaptive differential evolution algorithm · 2010
Earlier work this paper cites.
TEC (2010)
Das, S., Suganthan, P.N.: Differential evolution: A survey of the state-of-the-art · 2010
Earlier work this paper cites.
In: CEC (2013)
Padhye, N., Mittal, P., Deb, K.: Differential evolution: Performances and analyses · 2013
Earlier work this paper cites.
Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery (2014)
Bartz-Beielstein, T., Branke, J., Mehnen, J., Mersmann, O.: Evolutionary algorithms · 2014
Earlier work this paper cites.
arXiv preprint arXiv:1408.0101 (2014)
Kumar, S., Sharma, V.K., Kumari, R.: Memetic search in differential evolution algorithm · 2014
Earlier work this paper cites.
In: ECCV (2014)
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context · 2014
Earlier work this paper cites.
In: ICDSP (2014)
Shi, E.C., Leung, F.H., Law, B.N.: Differential evolution with adaptive population size · 2014
Earlier work this paper cites.
In: CVPR (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Earlier work this paper cites.
In: ICGTSPICC (2016)
Vikhar, P.A.: Evolutionary algorithms: A critical review and its future prospects · 2016
Earlier work this paper cites.
In: ICCV (2017)
Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., Wei, Y.: Deformable convolutional networks · 2017
Earlier work this paper cites.
In: ICCV (2017)
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad-cam: Visual explanations from deep networks via gradient-based localization · 2017
Earlier work this paper cites.
In: NeurIPS (2017)
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L.u., Polosukhin, I.: Attention is all you need · 2017
Earlier work this paper cites.
In: ICLR (2018)
Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., Zettlemoyer, L.: Deep contextualized word representations · 2018
Earlier work this paper cites.
OpenAI (2018)
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al.: Improving language understanding by generative pre-training · 2018
Earlier work this paper cites.
In: ECCV (2018)
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., Sun, J.: Unified perceptual parsing for scene understanding · 2018
Earlier work this paper cites.
arXiv preprint arXiv:1906.07155 (2019)
Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., Zhang, Z., Cheng, D., Zhu, C., Cheng, T., Zhao, Q., Li, B., Lu, X., Zhu, R., Wu, Y., Dai, J., Wang, J., Shi, J., Ouyang, W., Loy, C.C., Lin, D.: MMDetection: Open mmlab detection toolbox and benchmark · 2019
Earlier work this paper cites.
In: NAACL (2019)
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding · 2019
Earlier work this paper cites.
Information (2019)
Hassanat, A., Almohammadi, K., Alkafaween, E., Abunawas, E., Hammouri, A., Prasath, V.: Choosing mutation and crossover ratios for genetic algorithms—a review with a new dynamic approach · 2019
Earlier work this paper cites.
In: ICCV (2019)
Howard, A., Sandler, M., Chu, G., Chen, L.C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., et al.: Searching for mobilenetv3 · 2019
Earlier work this paper cites.
In: ICLR (2019)
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization · 2019
Earlier work this paper cites.
Swarm and evolutionary computation (2019)
Opara, K.R., Arabas, J.: Differential evolution: A survey of theoretical analyses · 2019
Earlier work this paper cites.
In: NeurIPS (2019)
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: Pytorch: An imperative style, high-performance deep learning library · 2019
Earlier work this paper cites.
OpenAI blog (2019)
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners · 2019
Earlier work this paper cites.
In: ICML (2019)
Tan, M., Le, Q.: Efficientnet: Rethinking model scaling for convolutional neural networks · 2019
Earlier work this paper cites.
https://github.com/rwightman/pytorch-image-models (2019)
Wightman, R.: Pytorch image models · 2019
Earlier work this paper cites.
IJCV (2019)
Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., Torralba, A.: Semantic understanding of scenes through the ade20k dataset · 2019
Earlier work this paper cites.
In: CVPR (2019)
Zhu, X., Hu, H., Lin, S., Dai, J.: Deformable convnets v2: More deformable, better results · 2019
Earlier work this paper cites.
In: NeurIPS (2020)
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners · 2020
Earlier work this paper cites.
In: ECCV (2020)
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers · 2020
Earlier work this paper cites.
In: ICLR (2020)
Cordonnier, J.B., Loukas, A., Jaggi, M.: On the relationship between self-attention and convolutional layers · 2020
Earlier work this paper cites.
arXiv preprint arXiv:2012.11747 (2020)
He, R., Ravula, A., Kanagal, B., Ainslie, J.: Realformer: Transformer likes residual attention · 2020
Earlier work this paper cites.
In: ICML (2020)
Katharopoulos, A., Vyas, A., Pappas, N., Fleuret, F.: Transformers are rnns: Fast autoregressive transformers with linear attention · 2020
Earlier work this paper cites.
In: ICLR (2020)
Kitaev, N., Kaiser, L., Levskaya, A.: Reformer: The efficient transformer · 2020
Earlier work this paper cites.
Engineering Applications of Artificial Intelligence (2020)
Pant, M., Zaheer, H., Garcia-Hernandez, L., Abraham, A., et al.: Differential evolution: A review of more than two decades of research · 2020
Earlier work this paper cites.
Genetic programming theory and practice XVII (2020)
Sloss, A.N., Gustafson, S.: 2019 evolutionary algorithms review · 2020
Earlier work this paper cites.
In: ACL (2020)
Wang, H., Wu, Z., Liu, Z., Cai, H., Zhu, L., Gan, C., Han, S.: Hat: Hardware-aware transformers for efficient natural language processing · 2020
Earlier work this paper cites.
arXiv preprint arXiv:2006.04768 (2020)
Wang, S., Li, B., Khabsa, M., Fang, H., Ma, H.: Linformer: Self-attention with linear complexity · 2020
Cited alongside, same era.
In: NeurIPS (2021)
Ali, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al.: Xcit: Cross-covariance image transformers · 2021
Cited alongside, same era.
arXiv preprint arXiv:2104.03602 (2021)
Atito, S., Awais, M., Kittler, J.: Sit: Self-supervised vision transformer · 2021
Cited alongside, same era.
In: ICLR (2021)
Bello, I.: Lambdanetworks: Modeling long-range interactions without attention · 2021
Cited alongside, same era.
In: ICML (2021)
Bertasius, G., Wang, H., Torresani, L.: Is space-time attention all you need for video understanding? · 2021
Cited alongside, same era.
In: ICCV (2021)
In: NeurIPS (2021)
Zhang, J., Xu, C., Li, J., Chen, W., Wang, Y., Tai, Y., Chen, S., Wang, C., Huang, F., Liu, Y.: Analogous to evolutionary algorithm: Designing a unified sequence model · 2021
Later among the works it cites.
In: CVPR (2021)
Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P.H., et al.: Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers · 2021
Later among the works it cites.
arXiv preprint arXiv:2103.11886 (2021)
Zhou, D., Kang, B., Jin, X., Yang, L., Lian, X., Hou, Q., Feng, J.: Deepvit: Towards deeper vision transformer · 2021
Later among the works it cites.
In: ICLR (2021)
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable {detr}: Deformable transformers for end-to-end object detection · 2021
Later among the works it cites.
In: ICML (2022)
Baevski, A., Hsu, W.N., Xu, Q., Babu, A., Gu, J., Auli, M.: Data2vec: A general framework for self-supervised learning in speech, vision and language · 2022
Closest in time.
In: ICLR (2022)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bhojanapalli, S., Chakrabarti, A., Glasner, D., Li, D., Unterthiner, T., Veit, A.: Understanding robustness of transformers for image classification · 2021
Cited alongside, same era.
J REAL-TIME IMAGE PR (2021)
Bhowmik, P., Pantho, M.J.H., Bobda, C.: Bio-inspired smart vision sensor: toward a reconfigurable hardware modeling of the hierarchical processing in the brain · 2021
Cited alongside, same era.
In: ICCV (2021)
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers · 2021
Cited alongside, same era.
In: ICCV (2021)
Chen, B., Li, P., Li, C., Li, B., Bai, L., Lin, C., Sun, M., Yan, J., Ouyang, W.: Glit: Neural architecture search for global and local image transformer · 2021
Cited alongside, same era.
In: CVPR (2021)
Chen, H., Wang, Y., Guo, T., Xu, C., Deng, Y., Liu, Z., Ma, S., Xu, C., Xu, C., Gao, W.: Pre-trained image processing transformer · 2021
Cited alongside, same era.
arXiv preprint arXiv:2102.04306 (2021)
Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, A.L., Zhou, Y.: Transunet: Transformers make strong encoders for medical image segmentation · 2021
Cited alongside, same era.
In: ICCV (2021)
Chen, M., Peng, H., Fu, J., Ling, H.: Autoformer: Searching transformers for visual recognition · 2021
Cited alongside, same era.
Bao, H., Dong, L., Piao, S., Wei, F.: BEit: BERT pre-training of image transformers · 2022
Closest in time.
In: CVPR (2022)
Chen, Q., Wu, Q., Wang, J., Hu, Q., Hu, T., Ding, E., Cheng, J., Wang, J.: Mixformer: Mixing features across windows and dimensions · 2022
Closest in time.
In: ICLR (2022)
Chen, T., Saxena, S., Li, L., Fleet, D.J., Hinton, G.: Pix2seq: A language modeling framework for object detection · 2022
Closest in time.
In: CVPR (2022)
Chen, Y., Dai, X., Chen, D., Liu, M., Dong, X., Yuan, L., Liu, Z.: Mobile-former: Bridging mobilenet and transformer · 2022
Closest in time.
In: CVPR (2022)
Cheng, B., Misra, I., Schwing, A.G., Kirillov, A., Girdhar, R.: Masked-attention mask transformer for universal image segmentation · 2022
Closest in time.
In: CVPR (2022)
Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., Chen, D., Guo, B.: Cswin transformer: A general vision transformer backbone with cross-shaped windows · 2022
Closest in time.
In: NeurIPS (2022)
Gao, P., Ma, T., Li, H., Lin, Z., Dai, J., Qiao, Y.: Mcmae: Masked convolution meets masked autoencoders · 2022
Closest in time.
Proceedings of the Royal Society A (2022)
Goyal, A., Bengio, Y.: Inductive biases for deep learning of higher-level cognition · 2022
Closest in time.
In: CVPR (2022)
Guo, J., Han, K., Wu, H., Tang, Y., Chen, X., Wang, Y., Xu, C.: Cmt: Convolutional neural networks meet vision transformers · 2022
Closest in time.
In: CVPR (2022)
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.: Masked autoencoders are scalable vision learners · 2022
Closest in time.
In: NeurIPS (2022)
Kim, J., Nguyen, D., Min, S., Cho, S., Lee, M., Lee, H., Hong, S.: Pure transformers are powerful graph learners · 2022
Closest in time.
In: CVPR (2022)
Lee, Y., Kim, J., Willette, J., Hwang, S.J.: Mpvit: Multi-path vision transformer for dense prediction · 2022
Closest in time.
In: ICLR (2022)
Li, K., Wang, Y., Peng, G., Song, G., Liu, Y., Li, H., Qiao, Y.: Uniformer: Unified transformer for efficient spatial-temporal representation learning · 2022
Closest in time.
In: ICML (2022)
Liu, Y., Li, H., Guo, Y., Kong, C., Li, J., Wang, S.: Rethinking attention-model explainability through faithfulness violation test · 2022
Closest in time.
In: CVPR (2022)
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al.: Swin transformer v2: Scaling up capacity and resolution · 2022
Closest in time.
In: ICLR (2022)
Mehta, S., Rastegari, M.: Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer · 2022
Closest in time.
In: NeurIPS (2022)
Min, J., Zhao, Y., Luo, C., Cho, M.: Peripheral vision transformer · 2022
Closest in time.
In: AAAI (2022)
Nakashima, K., Kataoka, H., Matsumoto, A., Iwata, K., Inoue, N., Satoh, Y.: Can vision transformers learn without natural images? · 2022
Closest in time.
In: CVPR (2022)
Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., Huang, G.: On the integration of self-attention and convolution · 2022
Closest in time.
In: NeurIPS (2022)
Qiang, Y., Pan, D., Li, C., Li, X., Jang, R., Zhu, D.: Attcat: Explaining transformers via attentive class activation tokens · 2022
Closest in time.
In: CVPR (2022)
Ren, S., Zhou, D., He, S., Feng, J., Wang, X.: Shunted self-attention via multi-scale token aggregation · 2022
Closest in time.
In: NeurIPS (2022)
Si, C., Yu, W., Zhou, P., Zhou, Y., Wang, X., YAN, S.: Inception transformer · 2022
Closest in time.
In: CVPR (2022)
Thatipelli, A., Narayan, S., Khan, S., Anwer, R.M., Khan, F.S., Ghanem, B.: Spatio-temporal relation modeling for few-shot action recognition · 2022
Closest in time.
In: ECCV (2022)
Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., Li, Y.: Maxvit: Multi-axis vision transformer · 2022
Closest in time.
In: CACV (2022)
Wan, Z., Chen, H., An, J., Jiang, W., Yao, C., Luo, J.: Facial attribute transformers for precise and robust makeup transfer · 2022
Closest in time.
In: CVPR (2022)
Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Jiang, Y.G., Zhou, L., Yuan, L.: Bevt: Bert pretraining of video transformers · 2022
Closest in time.
CVM (2022)
Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pvt v2: Improved baselines with pyramid vision transformer · 2022
Closest in time.
In: ICLR (2022)
Wang, W., Yao, L., Chen, L., Lin, B., Cai, D., He, X., Liu, W.: Crossformer: A versatile vision transformer hinging on cross-scale attention · 2022
Closest in time.
In: CVPR (2022)
Wei, C., Fan, H., Xie, S., Wu, C.Y., Yuille, A., Feichtenhofer, C.: Masked feature prediction for self-supervised visual pre-training · 2022
Closest in time.
In: CVPR (2022)
Xia, Z., Pan, X., Song, S., Li, L.E., Huang, G.: Vision transformer with deformable attention · 2022
Closest in time.
In: CVPR (2022)
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., Hu, H.: Simmim: A simple framework for masked image modeling · 2022
Closest in time.
In: CVPR (2022)
Yang, C., Wang, Y., Zhang, J., Zhang, H., Wei, Z., Lin, Z., Yuille, A.: Lite vision transformer with enhanced self-attention · 2022
Closest in time.
In: CVPR (2022)
Yu, W., Luo, M., Zhou, P., Si, C., Zhou, Y., Wang, X., Feng, J., Yan, S.: Metaformer is actually what you need for vision · 2022
Closest in time.
TPAMI (2022)
Yuan, L., Hou, Q., Jiang, Z., Feng, J., Yan, S.: Volo: Vision outlooker for visual recognition · 2022
Closest in time.
In: CVPR (2022)
Zamir, S.W., Arora, A., Khan, S., Hayat, M., Khan, F.S., Yang, M.H.: Restormer: Efficient transformer for high-resolution image restoration · 2022
Closest in time.
IJCV (2023)
Chen, X., Ding, M., Wang, X., Xin, Y., Mo, S., Wang, Y., Han, S., Luo, P., Zeng, G., Wang, J.: Context autoencoder for self-supervised representation learning · 2023
Closest in time.
In: ICLR (2023)
Chu, X., Tian, Z., Zhang, B., Wang, X., Wei, X., Xia, H., Shen, C.: Conditional positional encodings for vision transformers · 2023
Closest in time.
In: AAAI (2023)
Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., Yu, N., Guo, B.: Peco: Perceptual codebook for bert pre-training of vision transformers · 2023
Closest in time.
CVM (2023)
Guo, M.H., Lu, C.Z., Liu, Z.N., Cheng, M.M., Hu, S.M.: Visual attention network · 2023
Closest in time.
In: CVPR (2023)
Hassani, A., Walton, S., Li, J., Li, S., Shi, H.: Neighborhood attention transformer · 2023
Closest in time.
TPAMI (2023)
Li, K., Wang, Y., Zhang, J., Gao, P., Song, G., Liu, Y., Li, H., Qiao, Y.: Uniformer: Unifying convolution and self-attention for visual recognition · 2023
Closest in time.
In: ICCV (2023)
Li, Y., Hu, J., Wen, Y., Evangelidis, G., Salahi, K., Wang, Y., Tulyakov, S., Ren, J.: Rethinking vision transformers for mobilenet size and speed · 2023
Closest in time.
In: ECCVW (2023)
Maaz, M., Shaker, A., Cholakkal, H., Khan, S., Zamir, S.W., Anwer, R.M., Shahbaz Khan, F.: Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications · 2023
Closest in time.
JAIHC (2023)
Xu, L., Yan, X., Ding, W., Liu, Z.: Attribution rollout: a new way to interpret visual transformer · 2023
Closest in time.
In: ICCV (2023)
Zhang, J., Li, X., Li, J., Liu, L., Xue, Z., Zhang, B., Jiang, Z., Huang, T., Wang, Y., Wang, C.: Rethinking mobile block for efficient attention-based models · 2023
Closest in time.
IJCV (2023)
Zhang, Q., Xu, Y., Zhang, J., Tao, D.: Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond · 2023
Closest in time.