Fetching the paper…
Reading the bibliography…
Masked image modeling (MIM) has achieved promising results on various vision tasks.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Signature verification using a" siamese" time delay neural network
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah · 1993
Earlier work this paper cites.
Autoencoders, minimum description length and helmholtz free energy
G. E. Hinton and R. Zemel · 1993
Earlier work this paper cites.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
M. Cuturi · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
M. Caron, P. Bojanowski, A. Joulin, and M. Douze · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
The inaturalist species classification and detection dataset
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie · 2018
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin · 2018
Cited alongside, same era.
Contrastive multiview coding
Y. Tian, D. Krishnan, and P. Isola · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2021
Later among the works it cites.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
A. Bardes, J. Ponce, and Y. LeCun · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Later among the works it cites.
Exploring simple siamese representation learning
X. Chen and K. He · 2021
Later among the works it cites.
An empirical study of training self-supervised vision transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unified perceptual parsing for scene understanding
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Cited alongside, same era.
Learning representations by maximizing mutual information across views
P. Bachman, R. D. Hjelm, and W. Buchwalter · 2019
Cited alongside, same era.
Cascade r-cnn: high quality object detection and instance segmentation
Z. Cai and N. Vasconcelos · 2019
Cited alongside, same era.
Unsupervised pre-training of image features on non-curated data
M. Caron, P. Bojanowski, J. Mairal, and A. Joulin · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo · 2019
Cited alongside, same era.
X. Chen, S. Xie, and K. He · 2021
Later among the works it cites.
Peco: Perceptual codebook for bert pre-training of vision transformers
X. Dong, J. Bao, T. Zhang, D. Chen, W. Zhang, L. Yuan, D. Chen, F. Wen, and N. Yu · 2021
Later among the works it cites.
Benchmarking detection transfer learning with vision transformers
Y. Li, S. Xie, X. Chen, P. Dollar, K. He, and R. Girshick · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny · 2021
Later among the works it cites.
Masked siamese networks for label-efficient learning
M. Assran, M. Caron, I. Misra, P. Bojanowski, F. Bordes, P. Vincent, A. Joulin, M. Rabbat, and N. Ballas · 2022
Closest in time.
Context autoencoder for self-supervised representation learning
X. Chen, M. Ding, X. Wang, Y. Xin, S. Mo, Y. Wang, S. Han, P. Luo, G. Zeng, and J. Wang · 2022
Closest in time.
Corrupted image modeling for self-supervised visual pre-training
Y. Fang, L. Dong, H. Bao, X. Wang, and F. Wei · 2022
Closest in time.
Convmae: Masked convolution meets masked autoencoders
P. Gao, T. Ma, H. Li, J. Dai, and Y. Qiao · 2022
Closest in time.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Closest in time.
Architecture-agnostic masked image modeling–from vit back to cnn
S. Li, D. Wu, F. Wu, Z. Zang, K. Wang, L. Shang, B. Sun, H. Li, S. Li, et al · 2022
Closest in time.
The devil is in the frequency: Geminated gestalt autoencoder for self-supervised visual pre-training
H. Liu, X. Jiang, X. Li, A. Guo, D. Jiang, and B. Ren · 2022
Closest in time.
Extreme masking for learning instance and distributed visual representations
Z. Wu, Z. Lai, X. Sun, and S. Lin · 2022
Closest in time.