Fetching the paper…
Reading the bibliography…
Pre-training models on large scale datasets, like ImageNet, is a standard practice in computer vision.
Eleanor H. Rosch, “Natural categories,” Cognitive Psychology , 1973
1973
Earlier work this paper cites.
Geoffrey E Hinton and Ruslan R Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle, “Greedy layer-wise training of deep networks,” in Advances in neural information processing systems , 2007, pp. 153–160
2007
Earlier work this paper cites.
Marc Ranzato, Christopher Poultney, Sumit Chopra, Yann LeCun et al. , “Efficient learning of sparse representations with an energy-based model,” Advances in neural information processing systems , vol. 19, p. 1137, 2007
2007
Earlier work this paper cites.
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proceedings of the 25th international conference on Machine learning , 2008, pp. 1096–1103
2008
Earlier work this paper cites.
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition , 2009
2009
Earlier work this paper cites.
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, Pierre-Antoine Manzagol, and Léon Bottou, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion.” Journal of machine learning research , vol. 11, no. 12, 2010
2010
Earlier work this paper cites.
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei, “3d object representations for fine-grained categorization,” in IEEE Workshop on 3D Representation and Recognition (3dRR-13) , 2013
2013
Earlier work this paper cites.
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic, “Learning and transferring mid-level image representations using convolutional neural networks,” in Computer Vision and Pattern Recognition , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Alexey Dosovitskiy, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox, “Discriminative unsupervised feature learning with convolutional neural networks,” Advances in neural information processing systems , vol. 27, pp. 766–774, 2014
2014
Earlier work this paper cites.
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool, “Food-101 – mining discriminative components with random forests,” in European Conference on Computer Vision , 2014
2014
Earlier work this paper cites.
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision , 2014
2014
Earlier work this paper cites.
Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox, “Discriminative unsupervised feature learning with exemplar convolutional neural networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 38, no. 9, pp. 1734–1747, 2015
2015
Earlier work this paper cites.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” in Computer Vision and Pattern Recognition , 2016
2016
Earlier work this paper cites.
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros, “Context encoders: Feature learning by inpainting,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2536–2544
2016
Earlier work this paper cites.
Richard Zhang, Phillip Isola, and Alexei A Efros, “Colorful image colorization,” in European conference on computer vision . Springer, 2016, pp. 649–666
2016
Earlier work this paper cites.
Mehdi Noroozi and Paolo Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” in European conference on computer vision . Springer, 2016, pp. 69–84
2016
Earlier work this paper cites.
Junyuan Xie, Ross Girshick, and Ali Farhadi, “Unsupervised deep embedding for clustering analysis,” in International conference on machine learning . PMLR, 2016, pp. 478–487
2016
Earlier work this paper cites.
Jianwei Yang, Devi Parikh, and Dhruv Batra, “Joint unsupervised learning of deep representations and image clusters,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016
2016
Earlier work this paper cites.
Armand Joulin, Laurens Van Der Maaten, Allan Jabri, and Nicolas Vasilache, “Learning visual features from large weakly supervised data,” in European Conference on Computer Vision . Springer, 2016, pp. 67–84
2016
Earlier work this paper cites.
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick, “Mask r-cnn,” in International Conference on Computer Vision , 2017
2017
Earlier work this paper cites.
Joao Carreira and Andrew Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017
2017
Cited alongside, same era.
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba, “Scene parsing through ade20k dataset,” in Computer Vision and Pattern Recognition , 2017
2017
Cited alongside, same era.
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie, “Feature pyramid networks for object detection,” in Computer Vision and Pattern Recognition , 2017
2017
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever, “Generative pretraining from pixels,” in International Conference on Machine Learning , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie, “The inaturalist species classification and detection dataset,” Computer Vision and Pattern Recognition , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3733–3742
2018
Cited alongside, same era.
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze, “Deep clustering for unsupervised learning of visual features,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 132–149
2018
Cited alongside, same era.
Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten, “Exploring the limits of weakly supervised pretraining,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 181–196
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Kaiming He, Ross Girshick, and Piotr Dollár, “Rethinking im-agenet pre-training,” arXiv preprint arXiv: 1811.08883 , 2018
2018
Cited alongside, same era.
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun, “Unified perceptual parsing for scene understanding,” in European Conference on Computer Vision , 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, and Yutaka Satoh, “Pre-training without natural images,” in Proceedings of the Asian Conference on Computer Vision , 2020
2020
Later among the works it cites.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Xinlei Chen and Kaiming He, “Exploring simple siamese representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 15 750–15 758
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Yahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, and Marco Nadai, “Efficient training of visual transformers with small datasets,” in Advances in Neural Information Processing Systems , 2021
2021
Closest in time.
2021
Closest in time.