Fetching the paper…
Reading the bibliography…
Since the introduction of Vision Transformer (ViT), patchification has long been regarded as a de facto image tokenization approach for plain visual architectures.
A new approach to linear filtering and prediction problems
Kalman, R. E · 1960
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Caesar, H., Uijlings, J., and Ferrari, V · 2018
Earlier work this paper cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J · 2018
Earlier work this paper cites.
Cascade r-cnn: High quality object detection and instance segmentation
Cai, Z. and Vasconcelos, N · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., and Torralba, A · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Earlier work this paper cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Cited alongside, same era.
Revisiting resnets: Improved training and scaling strategies
Bello, I., Fedus, W., Du, X., Cubuk, E. D., Srinivas, A., Lin, T.-Y., Shlens, J., and Zoph, B · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., and Ré, C · 2021
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Later among the works it cites.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Efficientnetv2: Smaller models and faster training
Tan, M. and Le, Q · 2021
Cited alongside, same era.
Resnet strikes back: An improved training procedure in timm
Wightman, R., Touvron, H., and Jégou, H · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.-H., Tay, F. E., Feng, J., and Yan, S · 2021
Cited alongside, same era.
Deepvit: Towards deeper vision transformer
Zhou, D., Kang, B., Jin, X., Yang, L., Lian, X., Jiang, Z., Hou, Q., and Feng, J · 2021
Cited alongside, same era.
Later among the works it cites.
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Transformers are ssms: Generalized models and efficient algorithms through structured state space duality
Dao, T. and Gu, A · 2024
Later among the works it cites.
Mambavision: A hybrid mamba-transformer vision backbone
Hatamizadeh, A. and Kautz, J · 2024
Later among the works it cites.
Localmamba: Visual state space model with windowed selective scan
Huang, T., Pei, X., You, S., Wang, F., Qian, C., and Xu, C · 2024
Later among the works it cites.
Videomamba: State space model for efficient video understanding
Li, K., Li, X., Wang, Y., He, Y., Wang, Y., Wang, L., and Qiao, Y · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2024
Later among the works it cites.
Vmamba: Visual state space model
Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., and Liu, Y · 2024
Later among the works it cites.
An image is worth more than 16x16 patches: Exploring transformers on individual pixels
Nguyen, D.-K., Assran, M., Jain, U., Oswald, M. R., Snoek, C. G., and Chen, X · 2024
Later among the works it cites.
Plainmamba: Improving non-hierarchical mamba in visual recognition
Yang, C., Chen, Z., Espinosa, M., Ericsson, L., Wang, Z., Liu, J., and Crowley, E. J · 2024
Later among the works it cites.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., and Wang, X · 2024
Later among the works it cites.
Vit-linearizer: Distilling quadratic knowledge into linear-time vision models
Wei, G. and Chellappa, R · 2025
Closest in time.