Fetching the paper…
Reading the bibliography…
The immense success of deep learning based methods in computer vision heavily relies on large scale training datasets.
Hofmann T (2001) Unsupervised learning by probabilistic latent semantic analysis. Machine Learning
2001
Earlier work this paper cites.
Duygulu P, Barnard K, de Freitas JF, Forsyth DA (2002) Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary. In: ECCV
2002
Earlier work this paper cites.
Blei DM, Jordan MI (2003) Modeling annotated data. In: SIGIR
2003
Earlier work this paper cites.
Blei DM, Ng AY, Jordan MI (2003) Latent Dirichlet allocation. Journal of Machine Learning Research
2003
Earlier work this paper cites.
Sivic J, Russell BC, Efros AA, Zisserman A, Freeman WT (2005) Discovering objects and their location in images. In: ICCV
2005
Earlier work this paper cites.
Rosipal R, Krämer N (2006) Overview and recent advances in partial least squares. In: Subspace, latent structure and feature selection
2006
Earlier work this paper cites.
Huiskes MJ, Lew MS (2008) The MIR flickr retrieval evaluation. In: MIR
2008
Earlier work this paper cites.
Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: CVPR
2009
Earlier work this paper cites.
Kittur A, Chi EH, Suh B (2009) What’s in wikipedia?: mapping topics and conflict using socially annotated category structure. In: Proceedings of the SIGCHI conference on human factors in computing systems, ACM, pp 1509–1512
2009
Earlier work this paper cites.
Mobahi H, Collobert R, Weston J (2009) Deep learning from temporal coherence in video. In: ICML
2009
Earlier work this paper cites.
Everingham M, Van Gool L, Williams CK, Winn J, Zisserman A (2010) The pascal visual object classes (VOC) challenge. IJCV
2010
Earlier work this paper cites.
Feng Y, Lapata M (2010) Topic models for image annotation and text illustration. In: HLT
2010
Earlier work this paper cites.
Putthividhy D, Attias HT, Nagarajan SS (2010) Topic regression multi-modal latent Dirichlet allocation for image annotation. In: CVPR
2010
Earlier work this paper cites.
Rasiwasia N, Costa Pereira J, Coviello E, Doyle G, Lanckriet GR, Levy R, Vasconcelos N (2010) A new approach to cross-modal multimedia retrieval. In: ACM-MM
2010
Earlier work this paper cites.
Xiao J, Hays J, Ehinger KA, Oliva A, Torralba A (2010) Sun database: Large-scale scene recognition from abbey to zoo. In: CVPR
2010
Earlier work this paper cites.
Coates A, Lee H, Ng AY (2011) An analysis of single-layer networks in unsupervised feature learning. In: AISTATS
2011
Earlier work this paper cites.
Li A, Shan S, Chen X, Gao W (2011) Face recognition based on non-corresponding region matching. In: ICCV
2011
Earlier work this paper cites.
Ordonez V, Kulkarni G, Berg TL (2011) Im2text: Describing images using 1 million captioned photographs. In: NIPS
2011
Earlier work this paper cites.
Tsikrika T, Popescu A, Kludas J (2011) Overview of the Wikipedia image retrieval task at ImageCLEF 2011. In: CLEF (Notebook Papers/Labs/Workshop)
2011
Earlier work this paper cites.
Wang Y, Mori G (2011) Max-margin latent Dirichlet allocation for image classification and annotation. In: BMVC
2011
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: NIPS
2012
Earlier work this paper cites.
Sharma A, Kumar A, Daume H, Jacobs DW (2012) Generalized multiview analysis: A discriminative latent space. In: CVPR
2012
Cited alongside, same era.
Frome A, Corrado GS, Shlens J, Bengio S, Dean J, Mikolov T, et al. (2013) Devise: A deep visual-semantic embedding model. In: NIPS
2013
Cited alongside, same era.
Mikolov T, Chen K, Corrado G, Dean J (2013) Efficient estimation of word representations in vector space
2013
Cited alongside, same era.
Rasiwasia N, Vasconcelos N (2013) Latent Dirichlet allocation models for image classification. IEEE Transactions on Pattern Analysis and Machine Intelligence
2013
Cited alongside, same era.
Swersky K, Snoek J, Adams RP (2013) Multi-task bayesian optimization. In: NIPS
2013
Cited alongside, same era.
Krähenbühl P, Doersch C, Donahue J, Darrell T (2015) Data-dependent initializations of convolutional neural networks. In: ICLR
2015
Later among the works it cites.
Oquab M, Bottou L, Laptev I, Sivic J (2015) Is object localization for free?-weakly-supervised learning with convolutional neural networks. In: CVPR
2015
Later among the works it cites.
Paine TL, Khorrami P, Han W, Huang TS (2015) An analysis of unsupervised pre-training in light of recent advances. In: ICLR
2015
Later among the works it cites.
Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A, et al. (2015) Going deeper with convolutions. CVPR
2015
Later among the works it cites.
Wang X, Gupta A (2015) Unsupervised learning of visual representations using videos. In: CVPR
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang K, He R, Wang W, Wang L, Tan T (2013) Learning coupled feature spaces for cross-modal matching. In: ICCV
2013
Cited alongside, same era.
Dosovitskiy A, Springenberg JT, Riedmiller M, Brox T (2014) Discriminative unsupervised feature learning with convolutional neural networks. In: NIPS
2014
Cited alongside, same era.
Gong Y, Ke Q, Isard M, Lazebnik S (2014) A multi-view embedding space for modeling internet images, tags, and their semantics. International journal of computer vision
2014
Cited alongside, same era.
Jia Y, Shelhamer E, Donahue J, Karayev S, Long J, Girshick R, Guadarrama S, Darrell T (2014) Caffe: Convolutional architecture for fast feature embedding. In: ICM
2014
Cited alongside, same era.
Le Q, Mikolov T (2014) Distributed representations of sentences and documents. In: ICML
2014
Cited alongside, same era.
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: ECCV
2014
Cited alongside, same era.
Pennington J, Socher R, Manning C (2014) Glove: Global vectors for word representation. In: EMNLP
2014
Cited alongside, same era.
Yang S, Luo P, Loy CC, Shum KW, Tang X (2015) Deep representation learning with target coding. In: AAAI
2015
Later among the works it cites.
Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A (2015) Object detectors emerge in deep scene CNNs. In: NIPS
2015
Later among the works it cites.
Bilen H, Vedaldi A (2016) Weakly supervised deep detection networks. In: CVPR
2016
Later among the works it cites.
Dundar A, Jin J, Culurciello E (2016) Convolutional clustering for unsupervised learning. In: ICLR
2016
Later among the works it cites.
Joulin A, Grave E, Bojanowski P, Mikolov T (2016) Bag of tricks for efficient text classification. arXiv preprint arXiv:160701759
2016
Later among the works it cites.
Owens A, Wu J, McDermott JH, Freeman WT, Torralba A (2016) Ambient sound provides supervision for visual learning. In: ECCV
2016
Later among the works it cites.
Patel Y, Gomez L, Rusinol M, Karatzas D (2016) Dynamic lexicon generation for natural scene images. In: ECCV
2016
Later among the works it cites.
Pathak D, Krahenbuhl P, Donahue J, Darrell T, Efros AA (2016) Context encoders: Feature learning by inpainting. In: CVPR
2016
Later among the works it cites.
Wang D, Tan X (2016) Unsupervised feature learning with c-svddnet. Pattern Recognition
2016
Later among the works it cites.
Zhao J, Mathieu M, Goroshin R, Lecun Y (2016) Stacked what-where auto-encoders. In: ICLR
2016
Later among the works it cites.
Bojanowski P, Joulin A (2017) Unsupervised learning by predicting noise. ICML
2017
Later among the works it cites.
Gomez L, Patel Y, Rusiñol M, Karatzas D, Jawahar C (2017) Self-supervised learning of visual features through embedding images into text topic spaces. In: CVPR
2017
Later among the works it cites.
Gordo A, Larlus D (2017) Beyond instance-level image retrieval: Leveraging captions to learn a global visual representation for semantic retrieval. In: CVPR
2017
Later among the works it cites.
Salvador A, Hynes N, Aytar Y, Marin J, Ofli F, Weber I, Torralba A (2017) Learning cross-modal embeddings for cooking recipes and food images
2017
Later among the works it cites.
Wang X, He K, Gupta A (2017) Transitive invariance for self-supervised visual representation learning. In: ICCV
2017
Later among the works it cites.