Fetching the paper…
Reading the bibliography…
Recently, MLP-like vision models have achieved promising performances on mainstream visual recognition tasks.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Earlier work this paper cites.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2018
Earlier work this paper cites.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., and He, K · 2018
Earlier work this paper cites.
Local relation networks for image recognition
Hu, H., Zhang, Z., Xie, Z., and Lin, S · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
Vision permutator: A permutable mlp-like architecture for visual recognition
Hou, Q., Jiang, Z., Yuan, L., Cheng, M.-M., Yan, S., and Feng, J · 2021
Later among the works it cites.
All tokens matter: Token labeling for training better vision transformers
Jiang, Z.-H., Hou, Q., Yuan, L., Zhou, D., Shi, Y., Jin, X., Wang, A., and Feng, J · 2021
Later among the works it cites.
Transformers in vision: A survey
Khan, S., Naseer, M., Hayat, M., Zamir, S. W., Khan, F. S., and Shah, M · 2021
Later among the works it cites.
Convmlp: Hierarchical convolutional mlps for vision
Li, J., Hassani, A., Walton, S., and Shi, H · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., et al · 2020
Cited alongside, same era.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Cited alongside, same era.
Resnest: Split-attention networks
Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Lin, H., Zhang, Z., Sun, Y., He, T., Mueller, J., Manmatha, R., et al · 2020
Cited alongside, same era.
Random erasing data augmentation
Zhong, Z., Zheng, L., Kang, G., Li, S., and Yang, Y · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
Multiscale vision transformers
Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., and Feichtenhofer, C · 2021
Cited alongside, same era.
Revitalizing cnn attention via transformers in self-supervised visual representation learning
Ge, C., Liang, Y., Song, Y., Jiao, J., Wang, J., and Luo, P · 2021
Cited alongside, same era.
Lian, D., Yu, Z., Sun, X., and Gao, S · 2021
Later among the works it cites.
Global filter networks for image classification
Rao, Y., Zhao, W., Zhu, Z., Lu, J., and Zhou, J · 2021
Later among the works it cites.
Bottleneck transformers for visual recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Later among the works it cites.
Sparse mlp for image recognition: Is self-attention really necessary?
Tang, C., Zhao, Y., Wang, G., Luo, C., Xie, W., and Zeng, W · 2021
Later among the works it cites.
Raftmlp: Do mlp-based models dream of winning over computer vision?
Tatsunami, Y. and Taki, M · 2021
Later among the works it cites.
Synthesizer: Rethinking self-attention for transformer models
Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., and Zheng, C · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Later among the works it cites.
Improving vision transformers by revisiting high-frequency components
Bai, J., Yuan, L., Xia, S.-T., Yan, S., Li, Z., and Liu, W · 2022
Closest in time.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Liang, Y., Ge, C., Tong, Z., Song, Y., Wang, J., and Xie, P · 2022
Closest in time.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Closest in time.