Fetching the paper…
Reading the bibliography…
As the deep learning revolution marches on, self-supervised learning has garnered increasing attention in recent years thanks to its remarkable representation learning ability and the low dependence on labeled data.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2004
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
A. Coates, A. Ng, and H. Lee · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, E. Rahtu, J. Kannala, M. B. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang, T. A. Funkhouser, L. J. Guibas, P. Hanrahan, Q.-X. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, J. Xiao, L. Yi, and F. Yu · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. K. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
D. F. Campos, T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, L. Deng, and B. Mitra · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. K. Gupta · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al · 2016
Earlier work this paper cites.
Aid: A benchmark data set for performance evaluation of aerial scene classification
G.-S. Xia, J. Hu, F. Hu, B. Shi, X. Bai, Y. Zhong, L. Zhang, and X. Lu · 2016
Earlier work this paper cites.
Places: An image database for deep scene understanding
B. Zhou, A. Khosla, À. Lapedriza, A. Torralba, and A. Oliva · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, S. Vijayanarasimhan, C. Pantofaru, D. A. Ross, G. Toderici, Y. Li, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
The inaturalist challenge 2017 dataset
G. V. Horn, O. M. Aodha, Y. Song, A. Shepard, H. Adam, P. Perona, and S. J. Belongie · 2017
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, A. Natsev, M. Suleyman, and A. Zisserman · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. H. Hovy · 2017
Earlier work this paper cites.
Agedb: The first manually collected, in-the-wild age database
S. Moschoglou, A. Papaioannou, C. Sagonas, J. Deng, I. Kotsia, and S. Zafeiriou · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Earlier work this paper cites.
Gan dissection: Visualizing and understanding generative adversarial networks
D. Bau, J.-Y. Zhu, H. Strobelt, B. Zhou, J. B. Tenenbaum, W. T. Freeman, and A. Torralba · 2018
Earlier work this paper cites.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Y.-A. Chung and J. R. Glass · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance-level discrimination
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin · 2018
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
A. Baevski, S. Schneider, and M. Auli · 2019
Earlier work this paper cites.
A short note on the kinetics-700 human action dataset
J. Carreira, E. Noland, C. Hillier, and A. Zisserman · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollár, and R. B. Girshick · 2019
Earlier work this paper cites.
Strategies for pre-training graph neural networks
W. Hu, B. Liu, J. Gomes, M. Zitnik, P. Liang, V. Pande, and J. Leskovec · 2019
Earlier work this paper cites.
Contrastive predictive coding based feature for automatic speaker verification
C.-I. Lai · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Evaluating protein transfer learning with tape
R. Rao, N. Bhattacharya, N. Thomas, Y. Duan, P. Chen, J. Canny, P. Abbeel, and Y. Song · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2019
Earlier work this paper cites.
Smiles-bert: large scale unsupervised pre-training for molecular property prediction
S. Wang, Y. Guo, Y. Wang, H. Sun, and J. Huang · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, H. Zhou, A. rahman Mohamed, and M. Auli · 2020
Earlier work this paper cites.
Mam: Masked acoustic modeling for end-to-end speech-to-text translation
J. Chen, M. Ma, R. Zheng, and L. Huang · 2020
Earlier work this paper cites.
Generative pretraining from pixels
M. Chen, A. Radford, J. Wu, H. Jun, P. Dhariwal, D. Luan, and I. Sutskever · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, et al · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Gpt-gnn: Generative pre-training of graph neural networks
Z. Hu, Y. Dong, K. Wang, K.-W. Chang, and Y. Sun · 2020
Earlier work this paper cites.
Prototypical contrastive learning of unsupervised representations
J. Li, P. Zhou, C. Xiong, R. Socher, and S. C. H. Hoi · 2020
Earlier work this paper cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee · 2020
Earlier work this paper cites.
Self-supervised contrastive learning of protein representations by mutual information maximization
A. X. Lu, H. Zhang, M. Ghassemi, and A. Moses · 2020
Earlier work this paper cites.
Rareact: A video dataset of unusual interactions
A. Miech, J.-B. Alayrac, I. Laptev, J. Sivic, and A. Zisserman · 2020
Earlier work this paper cites.
Self-supervised graph transformer on large-scale molecular data
Y. Rong, Y. Bian, T. Xu, W. Xie, Y. Wei, W. Huang, and J. Huang · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn · 2020
Earlier work this paper cites.
Mp3: A unified model to map, perceive, predict and plan
S. Casas, A. Sadat, and R. Urtasun · 2021
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. K. Sharma, N. Ding, and R. Soricut · 2021
Earlier work this paper cites.
Audio albert: A lite bert for self-supervised learning of audio representation
P.-H. Chi, P.-H. Chung, T.-H. Wu, C.-C. Hsieh, Y.-H. Chen, S.-W. Li, and H.-y. Lee · 2021
Earlier work this paper cites.
Peco: Perceptual codebook for bert pre-training of vision transformers
X. Dong, J. Bao, T. Zhang, D. Chen, W. Zhang, L. Yuan, D. Chen, F. Wen, and N. Yu · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Are large-scale datasets necessary for self-supervised pre-training?
A. El-Nouby, G. Izacard, H. Touvron, I. Laptev, H. Jégou, and E. Grave · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis, 2021
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
T. Gao, X. Yao, and D. Chen · 2021
Earlier work this paper cites.
Pre-training co-evolutionary protein representation via a pairwise masked language model
L. He, S. Zhang, L. Wu, H. Xia, F. Ju, H. Zhang, S. Liu, Y. Xia, J. Zhu, P. Deng, et al · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, et al · 2021
Earlier work this paper cites.
Vision transformer for small-size datasets
S. H. Lee, S. Lee, and B. C. Song · 2021
Earlier work this paper cites.
Mst: Masked self-supervised transformer for visual representation
Z. Li, Z. Chen, F. Yang, W. Li, Y. Zhu, C. Zhao, R. Deng, L. Wu, R. Zhao, M. Tang, and J. Wang · 2021
Earlier work this paper cites.
Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d
Y. Liao, J. Xie, and A. Geiger · 2021
Earlier work this paper cites.
Self-supervised learning: Generative or contrastive
X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang · 2021
Earlier work this paper cites.
Adversarial contrastive pre-training for protein sequences
M. McDermott, B. Yap, H. Hsu, D. Jin, and P. Szolovits · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma, et al · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Earlier work this paper cites.
Masked feature prediction for self-supervised visual pre-training
C. Wei, H. Fan, S. Xie, C. Wu, A. L. Yuille, and C. Feichtenhofer · 2021
Earlier work this paper cites.
Simmim: a simple framework for masked image modeling
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Earlier work this paper cites.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Earlier work this paper cites.
Masked siamese networks for label-efficient learning
M. Assran, M. Caron, I. Misra, P. Bojanowski, F. Bordes, P. Vincent, A. Joulin, M. G. Rabbat, and N. Ballas · 2022
Earlier work this paper cites.
Mae-ast: Masked autoencoding audio spectrogram transformer
A. Baade, P. Peng, and D. F. Harwath · 2022
Earlier work this paper cites.
Multimae: Multi-modal multi-task masked autoencoders
R. Bachmann, D. Mizrahi, A. Atanov, and A. R. Zamir · 2022
Earlier work this paper cites.
Efficient self-supervised learning with contextualized target representations for vision, speech and language
A. Baevski, A. Babu, W.-N. Hsu, and M. Auli · 2022
Earlier work this paper cites.
data2vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli · 2022
Earlier work this paper cites.
Masked autoencoders enable efficient knowledge distillers
Y. Bai, Z. Wang, J. Xiao, C. Wei, H. Wang, A. L. Yuille, Y. Zhou, and C. Xie · 2022
Earlier work this paper cites.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2022
Earlier work this paper cites.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Cited alongside, same era.
Efficient self-supervised vision pretraining with local masked reconstruction
J. Chen, M. Hu, B. Li, and M. Elhoseiny · 2022
Cited alongside, same era.
Context autoencoder for self-supervised representation learning
X. Chen, M. Ding, X. Wang, Y. Xin, S. Mo, Y. Wang, S. Han, P. Luo, G. Zeng, and J. Wang · 2022
Cited alongside, same era.
Sdae: Self-distillated masked autoencoder
Y. Chen, Y. Liu, D. Jiang, X. Zhang, W. Dai, H. Xiong, and Q. Tian · 2022
Cited alongside, same era.
Mask-guided vision transformer (mg-vit) for few-shot learning
Y. Chen, Z. Xiao, L. Zhao, L. Zhang, H. Dai, D. Liu, Z. Wu, C. Li, T. Zhang, C. Li, D. Zhu, T. Liu, and X. Jiang · 2022
Cited alongside, same era.
Masked spectrogram prediction for self-supervised audio pre-training
D. Chong, H. Wang, P. Zhou, and Q. jie Zeng · 2022
Masked image modeling with denoising contrast
K. Yi, Y. Ge, X. Li, S. Yang, D. Li, J. Wu, Y. Shan, and X. Qie · 2022
Later among the works it cites.
Cross-modality and self-supervised protein embedding for compound–protein affinity and contact prediction
Y. You and Y. Shen · 2022
Later among the works it cites.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
X. Yu, L. Tang, Y. Rao, T. Huang, J. Zhou, and J. Lu · 2022
Later among the works it cites.
Position prediction as an effective pretraining strategy
S. Zhai, N. Jaitly, J. Ramapuram, D. Busbridge, T. Likhomanenko, J. Y. Cheng, W. A. Talbott, C. Huang, H. Goh, and J. M. Susskind · 2022
Later among the works it cites.
i-mae: Are latent representations in masked autoencoders linearly separable?
K. Zhang and Z. Shen · 2022
Later among the works it cites.
How mask matters: Towards theoretical understandings of masked autoencoders
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery
Y. Cong, S. Khanna, C. Meng, P. Liu, E. Rozi, Y. He, M. Burke, D. Lobell, and S. Ermon · 2022
Cited alongside, same era.
Bootstrapped masked autoencoders for vision bert pretraining
X. Dong, J. Bao, T. Zhang, D. Chen, W. Zhang, L. Yuan, D. Chen, F. Wen, and N. Yu · 2022
Cited alongside, same era.
Maskclip: Masked self-distillation advances contrastive language-image pretraining
X. Dong, Y. Zheng, J. Bao, T. Zhang, D. Chen, H. Yang, M. Zeng, W. Zhang, L. Yuan, D. Chen, F. Wen, and N. Yu · 2022
Cited alongside, same era.
Corrupted image modeling for self-supervised visual pre-training
Y. Fang, L. Dong, H. Bao, X. Wang, and F. Wei · 2022
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Y. Fang, W. Wang, B. Xie, Q.-S. Sun, L. Y. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao · 2022
Cited alongside, same era.
Unleashing vanilla vision transformer with masked image modeling for object detection
Y. Fang, S. Yang, S. Wang, Y. Ge, Y. Shan, and X. Wang · 2022
Cited alongside, same era.
Q. Zhang, Y. Wang, and Y. Wang · 2022
Later among the works it cites.
Point-m2ae: Multi-scale masked autoencoders for hierarchical point cloud pre-training
R. Zhang, Z. Guo, P. Gao, R. Fang, B. Zhao, D. Wang, Y. J. Qiao, and H. Li · 2022
Later among the works it cites.
Graph masked autoencoders with transformers
S. Zhang, H. Chen, H. Yang, X. Sun, P. S. Yu, and G. Xu · 2022
Later among the works it cites.
Cae v2: Context autoencoder with clip target
X. Zhang, J. Chen, J. Yuan, Q. Chen, J. Wang, X. Wang, S. Han, X. Chen, J. Pi, K. Yao, J. Han, E. Ding, and J. Wang · 2022
Later among the works it cites.
Integrally migrating pre-trained transformer encoder-decoders for visual object detection
X. Zhang, F. Liu, Z. Peng, Z. Guo, F. Wan, X.-W. Ji, and Q. Ye · 2022
Later among the works it cites.
Hivit: Hierarchical vision transformer meets masked image modeling
X. Zhang, Y. Tian, W. Huang, Q. Ye, Q. Dai, L. Xie, and Q. Tian · 2022
Later among the works it cites.
Cim: Constrained intrinsic motivation for sparse-reward continuous control
X. Zheng, X. Ma, and C. Wang · 2022
Later among the works it cites.
ibot: Image bert pre-training with online tokenizer
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong · 2022
Later among the works it cites.
Self pre-training with masked autoencoders for medical image analysis
L. Zhou, H. Liu, J. Bae, J. He, D. Samaras, and P. Prasanna · 2022
Later among the works it cites.
Self-supervised learning from images with a joint-embedding predictive architecture
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y. LeCun, and N. Ballas · 2023
Closest in time.
Sequential modeling enables scalable learning for large vision models, 2023
Y. Bai, X. Geng, K. Mangalam, A. Bar, A. Yuille, T. Darrell, J. Malik, and A. A. Efros · 2023
Closest in time.
Adamae: Adaptive masking for efficient spatiotemporal learning with masked autoencoders
W. G. C. Bandara, N. Patel, A. Gholami, M. Nikkhah, M. Agrawal, and V. M. Patel · 2023
Closest in time.
Learning to mask and permute visual tokens for vision transformer pre-training
L. Baraldi, R. Amoroso, M. Cornia, A. Pilzer, and R. Cucchiara · 2023
Closest in time.
Pimae: Point cloud and image interactive masked autoencoders for 3d object detection
A. Chen, K. Zhang, R. Zhang, Z. Wang, Y. Lu, Y. Guo, and S. Zhang · 2023
Closest in time.
Masked image training for generalizable deep image denoising
H. Chen, J. Gu, Y. Liu, S. A. Magid, C. Dong, Q. Wang, H. Pfister, and L. Zhu · 2023
Closest in time.
Traj-mae: Masked autoencoders for trajectory prediction
H. Chen, J. Wang, K. Shao, F. Liu, J. Hao, C. Guan, G. Chen, and P.-A. Heng · 2023
Closest in time.
Improving masked autoencoders by learning where to mask
H. Chen, W. Zhang, Y. Wang, and X. Yang · 2023
Closest in time.
Mixed autoencoder for self-supervised visual representation learning
K. Chen, Z. Liu, L. Hong, H. Xu, Z. Li, and D.-Y. Yeung · 2023
Closest in time.
Humanmac: Masked motion completion for human motion prediction
L. Chen, J. Zhang, Y. rong Li, Y. Pang, X. Xia, and T. Liu · 2023
Closest in time.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, Z. Muyan, Q. Zhang, X. Zhu, L. Lu, B. Li, P. Luo, T. Lu, Y. Qiao, and J. Dai · 2023
Closest in time.
Forecast-mae: Self-supervised pre-training for motion forecasting with masked autoencoders
J. Cheng, X. Mei, and M.-Y. Liu · 2023
Closest in time.
Autoencoders as cross-modal teachers: Can pretrained 2d image transformers help 3d representation learning?
R. Dong, Z. Qi, L. Zhang, J. Zhang, J. Sun, Z. Ge, L. Yi, and K. Ma · 2023
Closest in time.
Motion-guided masking for spatiotemporal representation learning
D. Fan, J. Wang, S. Liao, Y. Zhu, V. Bhat, H. J. Santos-Villalobos, M. V. Rohith, and X. Li · 2023
Closest in time.
Eva-02: A visual representation for neon genesis
Y. Fang, Q. Sun, X. Wang, T. Huang, X. Wang, and Y. Cao · 2023
Closest in time.
Instructcv: Instruction-tuned text-to-image diffusion models as vision generalists
Y. Gan, S. Park, A. Schubert, A. Philippakis, and A. M. Alaa · 2023
Closest in time.
Pre-training antibody language models for antigen-specific computational antibody design
K. Gao, L. Wu, J. Zhu, T. Peng, Y. Xia, L. He, S. Xie, T. Qin, H. Liu, K. He, et al · 2023
Closest in time.
Vqpl: Vector quantized protein language
Z. Gao, C. Tan, and S. Z. Li · 2023
Closest in time.
Instructdiffusion: A generalist modeling interface for vision tasks
Z. Geng, B. Yang, T. Hang, C. Li, S. Gu, T. Zhang, J. Bao, Z. Zhang, H. Hu, D. Chen, and B. Guo · 2023
Closest in time.
Audiovisual masked autoencoders
M.-I. Georgescu, E. Fonseca, R. T. Ionescu, M. Lucic, C. Schmid, and A. Arnab · 2023
Closest in time.
Siamese masked autoencoders
A. Gupta, J. Wu, J. Deng, and L. Fei-Fei · 2023
Closest in time.
Revcolv2: Exploring disentangled representations in masked image modeling
Q. Han, Y. Cai, and X. Zhang · 2023
Closest in time.
Graphmae2: A decoding-enhanced masked self-supervised graph learner
Z. Hou, Y. He, Y. Cen, X. Liu, Y. Dong, E. Kharlamov, and J. Tang · 2023
Closest in time.
Mgmae: Motion guided masking for video masked autoencoding
B. Huang, Z. Zhao, G. Zhang, Y. Qiao, and L. Wang · 2023
Closest in time.
Improving adversarial robustness of masked autoencoders via test-time frequency-domain prompting
Q. Huang, X. Dong, D. Chen, Y. Chen, L. Yuan, G. Hua, W. Zhang, N. H. Yu, and M. Reaserch · 2023
Closest in time.
Generic-to-specific distillation of masked autoencoders
W. Huang, Z. Peng, L. Dong, F. Wei, J. Jiao, and Q. Ye · 2023
Closest in time.
Layer grafted pre-training: Bridging contrastive learning and masked image modeling for label-efficient representations
Z. Jiang, Y. Chen, M. Liu, D. Chen, X. Dai, L. Yuan, Z. Liu, and Z. Wang · 2023
Closest in time.
Mesa: Masked, geometric, and supervised pre-training for monocular depth estimation
M. O. Khan, J. Liang, C.-K. Wang, S. Yang, and Y. Lou · 2023
Closest in time.
Understanding masked autoencoders via hierarchical latent variable models
L. Kong, M. Q. Ma, G. Chen, E. P. Xing, Y. Chi, L.-P. Morency, and K. Zhang · 2023
Closest in time.
Masked vision and language modeling for multi-modal representation learning
G. Kwon, Z. Cai, A. Ravichandran, E. Bas, R. Bhotika, and S. . Soatto · 2023
Closest in time.
Masked autoencoders are stronger knowledge distillers
S. Lao, G. Song, B. Liu, Y. Liu, and Y. Yang · 2023
Closest in time.
Contrastive tuning: A little help to make masked autoencoders forget
J. Lehner, B. Alkin, A. Fürst, E. Rumetshofer, L. Miklautz, and S. Hochreiter · 2023
Closest in time.
Dreamteacher: Pretraining image backbones with deep generative models
D. Li, H. Ling, A. Kar, D. Acuna, S. W. Kim, K. Kreis, A. Torralba, and S. Fidler · 2023
Closest in time.
Architecture-agnostic masked image modeling - from vit back to cnn
S. Li, D. Wu, F. Wu, Z. Zang, K. Wang, L. Shang, B. Sun, H. Li, and Stan.Z.Li · 2023
Closest in time.
Self-conditioned image generation via generating representations
T. Li, D. Katabi, and K. He · 2023
Closest in time.
Language quantized autoencoders: Towards unsupervised text-image alignment
H. Liu, W. Yan, and P. Abbeel · 2023
Closest in time.
Towards better 3d knowledge transfer via masked image modeling for multi-view 3d understanding
J. Liu, T. Wang, B. Liu, Q. Zhang, Y. Liu, and H. Li · 2023
Closest in time.
Docmae: Document image rectification via self-supervised representation learning
S. Liu, H. Feng, W. gang Zhou, H. Li, C. Liu, and F. Wu · 2023
Closest in time.
Pixmim: Rethinking pixel reconstruction in masked image modeling
Y. Liu, S. Zhang, J. Chen, K. Chen, and D. Lin · 2023
Closest in time.
Improving pixel-based mim by reducing wasted modeling capability
Y. Liu, S. Zhang, J. Chen, Z. Yu, K. Chen, and D. Lin · 2023
Closest in time.
Cmae-v: Contrastive masked autoencoders for video action recognition
C. Lu, X. Jin, Z. Huang, Q. Hou, M.-M. Cheng, and J. Feng · 2023
Closest in time.
Masked motion predictors are strong 3d action representation learners
Y. Mao, J. Deng, W. gang Zhou, Y. Fang, W. Ouyang, and H. Li · 2023
Closest in time.
Cmid: A unified self-supervised learning framework for remote sensing image understanding
D. Muhtar, X. liang Zhang, P. Xiao, Z. Li, and F. Gu · 2023
Closest in time.
R-mae: Regions meet masked autoencoders
D.-K. Nguyen, V. Aggarwal, Y. Li, M. R. Oswald, A. Kirillov, C. G. M. Snoek, and X. Chen · 2023
Closest in time.
Img2vec: A teacher of high token-diversity helps masked autoencoders, 2023
H. Pan, C. Liu, W. Wang, L. Yuan, H. Wang, Z. Li, and W. Liu · 2023
Closest in time.
Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining
Z. Qi, R. Dong, G. Fan, Z. Ge, X. Zhang, K. Ma, and L. Yi · 2023
Closest in time.
Rejuvenating image-gpt as strong visual representation learners, 2023
S. Ren, Z. Wang, H. Zhu, J. Xiao, A. Yuille, and C. Xie · 2023
Closest in time.
Tinymim: An empirical study of distilling mim pre-trained models
S. Ren, F. Wei, Z. Zhang, and H. Hu · 2023
Closest in time.
Hiera: A hierarchical vision transformer without the bells-and-whistles
C. K. Ryali, Y.-T. Hu, D. Bolya, C. Wei, H. Fan, P.-Y. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffman, J. Malik, Y. Li, and C. Feichtenhofer · 2023
Closest in time.
Pointcmp: Contrastive mask prediction for self-supervised learning on point cloud videos
Z. Shen, X. Sheng, L. Wang, Y. K. Guo, Q. Liu, and X. Zhou · 2023
Closest in time.
The effectiveness of mae pre-pretraining for billion-scale pretraining
M. Singh, Q. Duval, K. V. Alwala, H. Fan, V. Aggarwal, A. B. Adcock, A. Joulin, P. Doll’ar, C. Feichtenhofer, R. B. Girshick, R. Girdhar, and I. Misra · 2023
Closest in time.
Saprot: Protein language modeling with structure-aware vocabulary
J. Su, C. Han, Y. Zhou, J. Shan, X. Zhou, and F. Yuan · 2023
Closest in time.
Designing bert for convolutional networks: Sparse and hierarchical masked modeling
K. Tian, Y. Jiang, Q. Diao, C. Lin, L. Wang, and Z. Yuan · 2023
Closest in time.
Geomae: Masked geometric target prediction for self-supervised point cloud pre-training
X. Tian, H. Ran, Y. Wang, and H. Zhao · 2023
Closest in time.
Droppos: Pre-training vision transformers by reconstructing dropped positions
H. Wang, J. Fan, Y. Wang, K. Song, T. Wang, and Z. Zhang · 2023
Closest in time.
Hard patches mining for masked image modeling
H. Wang, K. Song, J. Fan, Y. Wang, J. Xie, and Z. Zhang · 2023
Closest in time.
Masked image modeling with local multi-scale reconstruction
H. Wang, Y. Tang, Y. Wang, J. Guo, Z. Deng, and K. Han · 2023
Closest in time.
Videomae v2: Scaling video masked autoencoders with dual masking
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao · 2023
Closest in time.
Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, L. Yuan, and Y.-G. Jiang · 2023
Closest in time.
Fremae: Fourier transform meets masked autoencoders for medical image segmentation
W. Wang, J. Wang, C. Chen, J. Jiao, L. Sun, Y. Cai, S. Song, and J. Li · 2023
Closest in time.
On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving, 2023
L. Wen, X. Yang, D. Fu, X. Wang, P. Cai, X. Li, T. Ma, Y. Li, L. Xu, D. Shang, Z. Zhu, S. Sun, Y. Bai, X. Cai, M. Dou, S. Hu, B. Shi, and Y. Qiao · 2023
Closest in time.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I.-S. Kweon, and S. Xie · 2023
Closest in time.
Dropmae: Masked autoencoders with spatial-attention dropout for tracking tasks
Q. Wu, T. Yang, Z. Liu, B. Wu, Y. Shan, and A. B. Chan · 2023
Closest in time.
Masked images are counterfactual samples for robust fine-tuning
Y. Xiao, Z. Tang, P. Wei, C. Liu, and L. Lin · 2023
Closest in time.
Skeletonmae: Graph-based masked autoencoder for skeleton sequence pre-training
H. Yan, Y. Liu, Y. Wei, Z. Li, G. Li, and L. Lin · 2023
Closest in time.
Unipad: A universal pre-training paradigm for autonomous driving
H. Yang, S. Zhang, D. Huang, X. Wu, H. Zhu, T. He, S. Tang, H. Zhao, Q. Qiu, B. Lin, X. He, and W. Ouyang · 2023
Closest in time.
Mrm: Masked relation modeling for medical image pre-training with genetics
Q. Yang, W. Li, B. Li, and Y. Yuan · 2023
Closest in time.
Moma: Distill from self-supervised teachers
Y. Yao, N. Desai, and M. S. Palaniswami · 2023
Closest in time.
Magvit: Masked generative video transformer
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M.-H. Yang, Y. Hao, I. Essa, and L. Jiang · 2023
Closest in time.
Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms
L. Yu, Y. Cheng, Z. Wang, V. Kumar, W. Macherey, Y. Huang, D. A. Ross, I. Essa, Y. Bisk, M. Yang, K. P. Murphy, A. G. Hauptmann, and L. Jiang · 2023
Closest in time.
Object recognition as next token prediction
K. Yue, B.-C. Chen, J. Geiping, H. Li, T. Goldstein, and S.-N. Lim · 2023
Closest in time.
Masked autoencoders are efficient class incremental learners
J.-T. Zhai, X. Liu, A. D. Bagdanov, K.-C. Li, and M.-M. Cheng · 2023
Closest in time.
A survey on masked autoencoder for self-supervised learning in vision and beyond
C. Zhang, C. Zhang, J. Song, J. S. K. Yi, K. Zhang, and I.-S. Kweon · 2023
Closest in time.
Learning 3d representations from 2d pre-trained models via image-to-point masked autoencoders
R. Zhang, L. Wang, Y. J. Qiao, P. Gao, and H. Li · 2023
Closest in time.
Contextual image masking modeling via synergized contrasting without view augmentation for faster and better visual pretraining
S. Zhang, F. Zhu, R. Zhao, and J. Yan · 2023
Closest in time.
Meta-transformer: A unified framework for multimodal learning
Y. Zhang, K. Gong, K. Zhang, H. Li, Y. J. Qiao, W. Ouyang, and X. Yue · 2023
Closest in time.
Protein representation learning by geometric structure pretraining
Z. Zhang, M. Xu, A. Jamasb, V. Chenthamarakshan, A. Lozano, P. Das, and J. Tang · 2023
Closest in time.
Masked retraining teacher-student framework for domain adaptive object detection
Z. Zhao, S. Wei, Q. Chen, D. Li, Y. Yang, Y. Peng, and Y. Liu · 2023
Closest in time.
Sparsemae: Sparse training meets masked autoencoders
A. Zhou, Y. Li, Z. Qin, J. Liu, J. Pan, R. Zhang, R. Zhao, P. Gao, and H. Li · 2023
Closest in time.
Masked autoencoders in computer vision: A comprehensive survey
Z. Zhou and X. Liu · 2023
Closest in time.
Vl-gpt: A generative pre-trained transformer for vision and language understanding and generation
J. Zhu, X. Ding, Y. Ge, Y. Ge, S. Zhao, H. Zhao, X. Wang, and Y. Shan · 2023
Closest in time.