Fetching the paper…
Reading the bibliography…
Recently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has dominated self-supervised learning in computer vision.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020, pp. 1877–1901
1901
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in CVPR , 2005, pp. 886–893 vol. 1
2005
Earlier work this paper cites.
Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2,” https://github.com/facebookresearch/detectron2
2007
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Master’s thesis, University of Tront , 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009, pp. 248–255
2009
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised visual representation learning by context prediction,” in ICCV , 2015, pp. 1422–1430
2015
Earlier work this paper cites.
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in CVPR , 2016, pp. 2536–2544
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 633–641
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Gupta, P. Dollar, and R. Girshick, “LVIS: A dataset for large vocabulary instance segmentation,” in CVPR , 2019, pp. 5356–5364
2019
Earlier work this paper cites.
S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shahbaz Khan, F. Zhu, L. Shao, G.-S. Xia, and X. Bai, “isaid: A large-scale dataset for instance segmentation in aerial images,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2019, pp. 28–37
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,” International Journal of Computer Vision , vol. 127, pp. 302–321, 2019
2019
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020, pp. 9729–9738
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML , 2020, pp. 1597–1607
2020
Earlier work this paper cites.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” in NeurIPS , 2020, pp. 21 271–21 284
2020
Earlier work this paper cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in ICML , 2020, pp. 1691–1703
2020
Cited alongside, same era.
O. J. Hénaff, A. Razavi, C. Doersch, S. Eslami, and A. v. d. Oord, “Data-efficient image recognition with contrastive predictive coding,” in ICML , 2020, pp. 4182–4192
2020
Cited alongside, same era.
H. Wang, X. Wu, Z. Huang, and E. P. Xing, “High-frequency component helps explain the generalization of convolutional neural networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 8684–8694
2020
Cited alongside, same era.
M. Contributors, “MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark,” https://github.com/open-mmlab/mmsegmentation
2020
Cited alongside, same era.
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in CVPR , June 2022, pp. 14 668–14 678
2022
Closest in time.
P. Wang, W. Zheng, T. Chen, and Z. Wang, “Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=O476oWmiNNp
2022
Closest in time.
X. He, L. Fang, M. Tan, and X. Chen, “Intra- and inter-slice contrastive learning for point supervised oct fluid segmentation,” IEEE Transactions on Image Processing , vol. 31, pp. 1870–1881, 2022
2022
Closest in time.
X. Wang, W. Wang, S. Yang, and J. Liu, “Clast: Contrastive learning for arbitrary style transfer,” IEEE Transactions on Image Processing , vol. 31, pp. 6761–6772, 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
N. Park and S. Kim, “How do vision transformers work?” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
P. Wang, W. Zheng, T. Chen, and Z. Wang, “Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
N. Zhao, Z. Wu, R. W. Lau, and S. Lin, “Distilling localization for self-supervised representation learning,” in AAAI , 2021, pp. 10 990–10 998
2021
Cited alongside, same era.
D. Dwibedi, Y. Aytar, J. Tompson, P. Sermanet, and A. Zisserman, “With a little help from my friends: Nearest-neighbor contrastive learning of visual representations,” in ICCV , 2021, pp. 9588–9597
2021
Cited alongside, same era.
R. R. Selvaraju, K. Desai, J. Johnson, and N. Naik, “Casting your model: Learning to localize improves self-supervised representations,” in CVPR , 2021, pp. 11 058–11 067
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , October 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
P. Chen, L. Li, J. Wu, W. Dong, and G. Shi, “Contrastive self-supervised pre-training for video quality assessment,” IEEE Transactions on Image Processing , vol. 31, pp. 458–471, 2022
2022
Closest in time.
X. Peng, K. Wang, Z. Zhu, M. Wang, and Y. You, “Crafting better contrastive views for siamese representation learning,” in CVPR , June 2022, pp. 16 031–16 040
2022
Closest in time.
M. Assefa, W. Jiang, K. Gedamu, G. Yilma, B. Kumeda, and M. Ayalew, “Self-supervised scene-debiasing for video representation learning via background patching,” IEEE Transactions on Multimedia , pp. 1–15, 2022
2022
Closest in time.
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo, “Swin transformer v2: Scaling up capacity and resolution,” in CVPR , June 2022, pp. 12 009–12 019
2022
Closest in time.
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong, “ibot: Image bert pre-training with online tokenizer,” in ICLR , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
N. Park and S. Kim, “How do vision transformers work?” in 10th International Conference on Learning Representations, ICLR 2022 , 2022
2022
Closest in time.
J. Bai, L. Yuan, S.-T. Xia, S. Yan, Z. Li, and W. Liu, “Improving vision transformers by revisiting high-frequency components,” in European Conference on Computer Vision , 2022
2022
Closest in time.
Z. Wang, H. Luo, P. WANG, F. Ding, F. Wang, and H. Li, “VTC-LFC: Vision transformer compression with low-frequency components,” in Thirty-Sixth Conference on Neural Information Processing Systems , 2022. [Online]. Available: https://openreview.net/forum?id=HuiLIB6EaOk
2022
Closest in time.
2022
Closest in time.
Z. Liu, J. Gui, and H. Luo, “Good helper is around you: Attention-driven masked image modeling,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 2, 2023, pp. 1799–1807
2023
Closest in time.
Y. Tian, Y. Yan, G. Zhai, L. Chen, and Z. Gao, “Clsa: A contrastive learning framework with selective aggregation for video rescaling,” IEEE Transactions on Image Processing , vol. 32, pp. 1300–1314, 2023
2023
Closest in time.
J. Wu, W. Sun, T. Gan, N. Ding, F. Jiang, J. Shen, and L. Nie, “Neighbor-guided consistent and contrastive learning for semi-supervised action recognition,” IEEE Transactions on Image Processing , vol. 32, pp. 2215–2227, 2023
2023
Closest in time.
2023
Closest in time.
Wikipedia contributors, “Inverse transform sampling — Wikipedia, the free encyclopedia,” 2024, [Online; accessed 20-August-2024]. [Online]. Available: https://en.wikipedia.org/wiki/Inverse_transform_sampling
2024
Closest in time.