Fetching the paper…
Reading the bibliography…
We propose ADIOS, a masked image model (MIM) framework for self-supervised learning, which simultaneously learns a masking function and an image encoder using an adversarial objective.
Signature verification using a “siamese” time delay neural network
Bromley, J., Bentz, J. W., Bottou, L., Guyon, I., Lecun, Y., Moore, C., Säckinger, E., and Shah, R · 1993
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R. B., and He, K · 2003
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Krähenbühl, P. and Koltun, V · 2011
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Pathak, D., Krähenbühl, P., Donahue, J., Darrell, T., and Efros, A. A · 2016
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B · 2017
Earlier work this paper cites.
The inaturalist species classification and detection dataset
Horn, G. V., Aodha, O. M., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S. J · 2018
Earlier work this paper cites.
Autoaugment: Learning augmentation strategies from data
Cubuk, E. D., Zoph, B., Mané, D., Vasudevan, V., and Le, Q. V · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Multi-object representation learning with iterative variational inference
Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E · 2020
Cited alongside, same era.
GENESIS: generative scene inference and sampling with object-centric latent representations
Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I · 2020
Cited alongside, same era.
Scan: Learning to classify images without labels
Gansbeke, W. V., Vandenhende, S., Georgoulis, S., Proesmans, M., and Gool, L. V · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Genesis-v2: Inferring unordered object representations without iterative refinement
Engelcke, M., Jones, O. P., and Posner, I · 2021
Later among the works it cites.
Whitening for self-supervised representation learning
Ermolov, A., Siarohin, A., Sangineto, E., and Sebe, N · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
Contrastive representation learning with trainable augmentation channel
Koyama, M., Minami, K., Miyato, T., and Gal, Y · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grill, J., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. Á., Guo, Z., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M · 2020
Cited alongside, same era.
Faster autoaugment: Learning augmentation strategies using backpropagation
Hataya, R., Zdenek, J., Yoshizoe, K., and Nakayama, H · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. B · 2020
Cited alongside, same era.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2021
Cited alongside, same era.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Bardes, A., Ponce, J., and LeCun, Y · 2021
Cited alongside, same era.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Cited alongside, same era.
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Viewmaker networks: Learning views for unsupervised representation learning
Tamkin, A., Wu, M., and Goodman, N. D · 2021
Later among the works it cites.
Augmenting convolutional networks with attention-based aggregation
Touvron, H., Cord, M., El-Nouby, A., Bojanowski, P., Joulin, A., Synnaeve, G., and Jégou, H · 2021
Later among the works it cites.
Noise or signal: The role of image backgrounds in object recognition
Xiao, K. Y., Engstrom, L., Ilyas, A., and Madry, A · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.
ibot: Image bert pre-training with online tokenizer
Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T · 2021
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M · 2022
Closest in time.
A convnet for the 2020s
Liu, Z., Mao, H., Chao-Yuan, W., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Closest in time.