Fetching the paper…
Reading the bibliography…
We present an efficient approach for Masked Image Modeling (MIM) with hierarchical Vision Transformers (ViTs), allowing the hierarchical ViTs to discard masked patches and operate only on the visible ones.
Dynamic programming
R. Bellman · 1966
Earlier work this paper cites.
Knapsack problems
H. Kellerer, U. Pferschy, and D. Pisinger · 2004
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Sparse 3d convolutional neural networks
B. Graham · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Earlier work this paper cites.
Fairness in machine learning
S. Barocas, M. Hardt, and A. Narayanan · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Colorization as a proxy task for visual understanding
G. Larsson, M. Maire, and G. Shakhnarovich · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
S. Gidaris, P. Singh, and N. Komodakis · 2018
Earlier work this paper cites.
3d semantic segmentation with submanifold sparse convolutional networks
B. Graham, M. Engelcke, and L. Van Der Maaten · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin · 2018
Cited alongside, same era.
Mmdetection: Open mmlab detection toolbox and benchmark
K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, et al · 2019
Cited alongside, same era.
4d spatio-temporal convnets: Minkowski convolutional neural networks
C. Choy, J. Gwak, and S. Savarese · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Interlaced sparse self-attention for semantic segmentation
L. Huang, Y. Yuan, J. Guo, C. Zhang, X. Chen, and J. Wang · 2019
Cited alongside, same era.
Self-adaptive training: Bridging supervised and self-supervised learning
L. Huang, C. Zhang, and H. Zhang · 2021
Later among the works it cites.
Shuffle transformer: Rethinking spatial shuffle for vision transformer
Z. Huang, Y. Ben, G. Luo, P. Cheng, G. Yu, and B. Fu · 2021
Later among the works it cites.
Benchmarking detection transfer learning with vision transformers
Y. Li, S. Xie, X. Chen, P. Dollar, K. He, and R. Girshick · 2021
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, et al · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Exploring simple Siamese representation learning
X. Chen and K. He · 2020
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Dense contrastive learning for self-supervised visual pre-training
X. Wang, R. Zhang, C. Shen, T. Kong, and L. Li · 2021
Later among the works it cites.
Masked feature prediction for self-supervised visual pre-training
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang · 2021
Later among the works it cites.
Region similarity representation learning
T. Xiao, C. J. Reed, X. Wang, K. Keutzer, and T. Darrell · 2021
Later among the works it cites.
Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning
Z. Xie, Y. Lin, Z. Zhang, Y. Cao, S. Lin, and H. Hu · 2021
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2021
Later among the works it cites.
A survey on green deep learning
J. Xu, W. Zhou, Z. Fu, H. Zhou, and L. Li · 2021
Later among the works it cites.
Adavit: Adaptive tokens for efficient vision transformer
H. Yin, A. Vahdat, J. Alvarez, A. Mallya, J. Kautz, and P. Molchanov · 2021
Later among the works it cites.
Hrformer: High-resolution vision transformer for dense predict
Y. Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny · 2021
Later among the works it cites.
Weakly supervised contrastive learning
M. Zheng, F. Wang, S. You, C. Qian, C. Zhang, X. Wang, and C. Xu · 2021
Later among the works it cites.
Ressl: Relational self-supervised learning with weak augmentation
M. Zheng, S. You, F. Wang, C. Qian, C. Zhang, X. Wang, and C. Xu · 2021
Later among the works it cites.
ibot: Image bert pre-training with online tokenizer
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong · 2021
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli · 2022
Closest in time.
Cmt: Convolutional neural networks meet vision transformers
J. Guo, K. Han, H. Wu, Y. Tang, X. Chen, Y. Wang, and C. Xu · 2022
Closest in time.
Learning where to learn in cross-view self-supervised learning
L. Huang, S. You, M. Zheng, F. Wang, C. Qian, and T. Yamasaki · 2022
Closest in time.
Metaformer is actually what you need for vision
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan · 2022
Closest in time.