Fetching the paper…
Reading the bibliography…
Masked image modeling (MIM) learns representations with remarkably good fine-tuning performances, overshadowing previous prevalent pre-training approaches such as image classification, instance contrastive learning, and image-text alignment.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. and Salakhutdinov, R. (2006) · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. (2020) · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2013) · 2013
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T. (2014) · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T. (2015) · 2015
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q. (2016) · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M. and Favaro, P. (2016) · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A. (2016) · 2016
Earlier work this paper cites.
Colorful image colorization
Zhang, R., Isola, P., and Efros, A. A. (2016) · 2016
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2017) · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B. (2017) · 2017
Cited alongside, same era.
Split-brain autoencoders: Unsupervised learning by cross-channel prediction
Zhang, R., Isola, P., and Efros, A. A. (2017) · 2017
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M. (2018) · 2018
Cited alongside, same era.
Unsupervised representation learning by predicting image rotations
Gidaris, S., Singh, P., and Komodakis, N. (2018) · 2018
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D. (2018) · 2018
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021) · 2021
Later among the works it cites.
Divide and contrast: Self-supervised learning from uncurated data
Tian, Y., Hénaff, O. J., and van den Oord, A. (2021) · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. (2021) · 2021
Later among the works it cites.
Dense contrastive learning for self-supervised visual pre-training
Wang, X., Zhang, R., Shen, C., Kong, T., and Li, L. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unified perceptual parsing for scene understanding
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J. (2018) · 2018
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al. (2020) · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F. (2021) · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021) · 2021
Cited alongside, same era.
An empirical study of training self-supervised vision transformers
Chen, X., Xie, S., and He, K. (2021) · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Cited alongside, same era.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L. (2021) · 2021
Later among the works it cites.
A simple baseline for zero-shot semantic segmentation with pre-trained vision-language model
Xu, M., Zhang, Z., Wei, F., Lin, Y., Cao, Y., Hu, H., and Bai, X. (2021) · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al. (2021) · 2021
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. (2021) · 2021
Later among the works it cites.
Deepvit: Towards deeper vision transformer
Zhou, D., Kang, B., Jin, X., Yang, L., Lian, X., Hou, Q., and Feng, J. (2021) · 2021
Later among the works it cites.
Vision transformer adapter for dense predictions
Chen, Z., Duan, Y., Wang, W., He, J., Lu, T., Dai, J., and Qiao, Y. (2022) · 2022
Closest in time.
icar: Bridging image classification and image-text alignment for visual recognition
Wei, Y., Cao, Y., Zhang, Z., Yao, Z., Xie, Z., Hu, H., and Guo, B. (2022) · 2022
Closest in time.
Simmim: A simple framework for masked image modeling
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. (2022) · 2022
Closest in time.
Unified contrastive learning in image-text-label space
Yang, J., Li, C., Zhang, P., Xiao, B., Liu, C., Yuan, L., and Gao, J. (2022) · 2022
Closest in time.