Fetching the paper…
Reading the bibliography…
Joint-embedding based learning (e.g., SimCLR, MoCo, DINO) and reconstruction-based learning (e.g., BEiT, SimMIM, MAE) are the two leading paradigms for self-supervised learning of vision transformers, but they differ substantially in their transfer performance.
A kernel statistical test of independence
A. Gretton, K. Fukumizu, C. Teo, L. Song, B. Schölkopf, and A. Smola · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2017
G. Alain and Y. Bengio · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
A. Morcos, M. Raghu, and S. Bengio · 2018
Earlier work this paper cites.
Similarity of neural network representations revisited
S. Kornblith, M. Norouzi, H. Lee, and G. Hinton · 2019
Earlier work this paper cites.
A critical analysis of self-supervision, or what we can learn from a single image
Y. Asano, C. Rupprecht, and A. Vedaldi · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. E. Hinton · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Exploring simple siamese representation learning
X. Chen and K. He · 2020
Cited alongside, same era.
Bootstrap your own latent: A new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Do self-supervised and supervised methods learn similar visual representations?
T. G. Grigg, D. Busbridge, J. Ramapuram, and R. Webb · 2021
Later among the works it cites.
Intriguing properties of vision transformers
M. M. Naseer, K. Ranasinghe, S. H. Khan, M. Hayat, F. Shahbaz Khan, and M.-H. Yang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Masked siamese networks for label-efficient learning
M. Assran, M. Caron, I. Misra, P. Bojanowski, F. Bordes, P. Vincent, A. Joulin, M. Rabbat, and N. Ballas · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Nguyen, M. Raghu, and S. Kornblith · 2020
Cited alongside, same era.
What is being transferred in transfer learning?
B. Neyshabur, H. Sedghi, and C. Zhang · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
Image bert pre-training with online tokenizer
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong · 2021
Cited alongside, same era.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
A. Bardes, J. Ponce, and Y. LeCun · 2021
Cited alongside, same era.
Understanding dimensional collapse in contrastive self-supervised learning
L. Jing, P. Vincent, Y. LeCun, and Y. Tian · 2021
Cited alongside, same era.
Are large-scale datasets necessary for self-supervised pre-training?
A. El-Nouby, G. Izacard, H. Touvron, I. Laptev, H. Jegou, and E. Grave · 2021
Cited alongside, same era.
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2022
Later among the works it cites.
High fidelity visualization of what your self-supervised representation knows about
F. Bordes, R. Balestriero, and P. Vincent · 2022
Later among the works it cites.
Head2Toe: Utilizing intermediate representations for better transfer learning
U. Evci, V. Dumoulin, H. Larochelle, and M. C. Mozer · 2022
Later among the works it cites.
A study of face obfuscation in imagenet
K. Yang, J. Yau, L. Fei-Fei, J. Deng, and O. Russakovsky · 2022
Later among the works it cites.
Exploring plain vision transformer backbones for object detection
Y. Li, H. Mao, R. Girshick, and K. He · 2022
Later among the works it cites.
The hidden uniform cluster prior in self-supervised learning
M. Assran, R. Balestriero, Q. Duval, F. Bordes, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, and N. Ballas · 2022
Later among the works it cites.
Omnimae: Single model masked pretraining on images and videos
R. Girdhar, A. El-Nouby, M. Singh, K. V. Alwala, A. Joulin, and I. Misra · 2022
Later among the works it cites.
What do self-supervised vision transformers learn?
N. Park, W. Kim, B. Heo, T. Kim, and S. Yun · 2023
Closest in time.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, and S. Xie · 2023
Closest in time.