Fetching the paper…
Reading the bibliography…
This paper enhances image-GPT (iGPT), one of the pioneering works that introduce autoregressive pretraining to predict the next pixels for visual representation learning.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K · 2003
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Pixel recurrent neural networks
Van Den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A. and Narasimhan, K · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
Wang, H., Ge, S., Lipton, Z., and Xing, E. P · 2019
Earlier work this paper cites.
XLNet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J. G., Salakhutdinov, R., and Le, Q. V · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ADE20K dataset
Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., and Torralba, A · 2019
Earlier work this paper cites.
Beyer, L., Hénaff, O. J., Kolesnikov, A., Zhai, X., and Oord, A. v. d · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T. J., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Coatnet: Marrying convolution and attention for all data sizes
Dai, Z., Liu, H., Le, Q. V., and Tan, M · 2021
Cited alongside, same era.
Peco: Perceptual codebook for bert pre-training of vision transformers
Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., and Yu, N · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Maxvit: Multi-axis vision transformer
Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., and Li, Y · 2022
Later among the works it cites.
Image as a foreign language: BEiT pretraining for all vision and vision-language tasks
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al · 2022
Later among the works it cites.
Contrastive learning rivals masked image modeling in fine-tuning via feature distillation
Wei, Y., Hu, H., Xie, Z., Zhang, Z., Cao, Y., Bao, J., Chen, D., and Guo, B · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tokenlearner: What can 8 learned tokens do for images and videos?
Ryoo, M. S., Piergiovanni, A., Arnab, A., Dehghani, M., and Angelova, A · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Masked feature prediction for self-supervised visual pre-training
Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., and Feichtenhofer, C · 2021
Cited alongside, same era.
ibot: Image bert pre-training with online tokenizer
Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T · 2021
Cited alongside, same era.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M · 2022
Cited alongside, same era.
BEiT: BERT pre-training of image transformers
Bao, H., Dong, L., Piao, S., and Wei, F · 2022
Cited alongside, same era.
Sdae: Self-distillated masked autoencoder
Chen, Y., Liu, Y., Jiang, D., Zhang, X., Dai, W., Xiong, H., and Tian, Q · 2022
Cited alongside, same era.
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., et al · 2023
Closest in time.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Closest in time.
An inverse scaling law for clip training
Li, X., Wang, Z., and Xie, C · 2023
Closest in time.
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Exploring stochastic autoregressive image modeling for visual representation
Qi, Y., Yang, F., Zhu, Y., Liu, Y., Wu, L., Zhao, R., and Li, W · 2023
Closest in time.
Imagenet-hard: The hardest images remaining from a study of the power of zoom and spatial biases in image classification
Taesiri, M. R., Nguyen, G., Habchi, S., Bezemer, C.-P., and Nguyen, A · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
One-peace: Exploring one general representation model toward unlimited modalities
Wang, P., Wang, S., Lin, J., Bai, S., Zhou, X., Zhou, J., Wang, X., and Zhou, C · 2023
Closest in time.
Scalable pre-training of large autoregressive image models
El-Nouby, A., Klein, M., Zhai, S., Bautista, M. A., Toshev, A., Shankar, V., Susskind, J. M., and Joulin, A · 2024
Closest in time.