Fetching the paper…
Reading the bibliography…
Recent advances in self-supervised learning (SSL) have largely closed the gap with supervised ImageNet pretraining.
Neil Bruce and John Tsotsos, “Saliency based on information maximization,”
2005
Earlier work this paper cites.
Raia Hadsell, Sumit Chopra, and Yann LeCun, “Dimensionality reduction by learning an invariant mapping,” in
2006
Earlier work this paper cites.
Jonathan Harel, Christof Koch, and Pietro Perona, “Graph-based visual saliency,” in
2007
Earlier work this paper cites.
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman, “The pascal visual object classes (VOC) challenge,”
2009
Earlier work this paper cites.
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
Michael Gutmann and Aapo Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in
2010
Earlier work this paper cites.
Stas Goferman, Lihi Zelnik-Manor, and Ayellet Tal, “Context-aware saliency detection,”
2011
Earlier work this paper cites.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
Huaizu Jiang, Jingdong Wang, Zejian Yuan, Yang Wu, Nanning Zheng, and Shipeng Li, “Salient object detection: A discriminative regional feature integration approach,” in
2013
Earlier work this paper cites.
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in
2014
Earlier work this paper cites.
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, “Microsoft COCO: Common objects in context,” in
2014
Earlier work this paper cites.
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva, “Learning deep features for scene recognition using places database,” in
2014
Earlier work this paper cites.
Wangjiang Zhu, Shuang Liang, Yichen Wei, and Jian Sun, “Saliency optimization from robust background detection,” in
2014
Earlier work this paper cites.
Ming-Ming Cheng, Niloy J Mitra, Xiaolei Huang, Philip HS Torr, and Shi-Min Hu, “Global contrast based salient region detection,”
2014
Earlier work this paper cites.
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein,
2015
Earlier work this paper cites.
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Earlier work this paper cites.
Jonathan Long, Evan Shelhamer, and Trevor Darrell, “Fully convolutional networks for semantic segmentation,” in
2015
Earlier work this paper cites.
Carl Doersch, Abhinav Gupta, and Alexei A. Efros, “Unsupervised visual representation learning by context prediction,” in
2015
Earlier work this paper cites.
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh, “VQA: Visual question answering,” in
2015
Earlier work this paper cites.
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li, “YFCC100M: The new data in multimedia research,”
2016
Earlier work this paper cites.
Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros, “Context encoders: Feature learning by inpainting,” in
2016
Earlier work this paper cites.
Richard Zhang, Phillip Isola, and Alexei A. Efros, “Colorful image colorization,” in
2016
Earlier work this paper cites.
Mehdi Noroozi and Paolo Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick, “Mask R-CNN,” in
2017
Cited alongside, same era.
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta, “Revisiting unreasonable effectiveness of data in deep learning era,” in
2017
Cited alongside, same era.
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in
2017
Cited alongside, same era.
Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, and Devi Parikh, “Taking a hint: Leveraging explanations to make vision and language models more grounded,” in
2019
Later among the works it cites.
Tam Nguyen, Maximilian Dax, Chaithanya Kumar Mummadi, Nhung Ngo, Thi Hoai Phuong Nguyen, Zhongyu Lou, and Thomas Brox, “Deepusps: Deep robust unsupervised saliency prediction via self-supervision,” in
2019
Later among the works it cites.
Priya Goyal, Dhruv Mahajan, Abhinav Gupta, and Ishan Misra, “Scaling and benchmarking self-supervised visual representation learning,” in
2019
Later among the works it cites.
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick, “Detectron2.”
2019
Later among the works it cites.
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, “A simple framework for contrastive learning of visual representations,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard Zhang, Phillip Isola, and Alexei A. Efros, “Split-brain autoencoders: Unsupervised learning by cross-channel prediction,” in
2017
Cited alongside, same era.
Dingwen Zhang, Junwei Han, and Yu Zhang, “Supervision by fusion: Towards unsupervised learning of deep salient object detector,” in
2017
Cited alongside, same era.
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie, “Feature pyramid networks for object detection,” in
2017
Cited alongside, same era.
Zhirong Wu, Yuanjun Xiong, Stella Yu, and Dahua Lin, “Unsupervised feature learning via non-parametric instance-level discrimination,” 2018
2018
Cited alongside, same era.
Dhruv Kumar Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten, “Exploring the limits of weakly supervised pretraining,” in
2018
Cited alongside, same era.
Spyros Gidaris, Praveer Singh, and Nikos Komodakis, “Unsupervised representation learning by predicting image rotations,” in
2018
Cited alongside, same era.
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze, “Deep clustering for unsupervised learning of visual features,” in
2018
Cited alongside, same era.
2020
Closest in time.
Ishan Misra and Laurens van der Maaten, “Self-supervised learning of pretext-invariant representations,” in
2020
Closest in time.
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, “Momentum contrast for unsupervised visual representation learning,” in
2020
Closest in time.
2020
Closest in time.
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin, “Unsupervised learning of visual features by contrasting cluster assignments,” in
2020
Closest in time.
Karan Desai and Justin Johnson, “Virtex: Learning visual representations from textual annotations,”
2020
Closest in time.
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar,
2020
Closest in time.
2020
Closest in time.
Yonglong Tian, Dilip Krishnan, and Phillip Isola, “Contrastive multiview coding,” in
2020
Closest in time.
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton, “Big self-supervised models are strong semi-supervised learners,”
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Xiao Zhang and Michael Maire, “Self-supervised visual representation learning from hierarchical grouping,”
2020
Closest in time.
Chih-Yao Ma, Yannis Kalantidis, Ghassan AlRegib, Peter Vajda, Marcus Rohrbach, and Zsolt Kira, “Learning to generate grounded visual captions without localization supervision,” in
2020
Closest in time.