Fetching the paper…
Reading the bibliography…
Is strong supervision necessary for learning a good visual representation? Do we really need millions of semantically-labeled images to train a Convolutional Neural Network (CNN)? In this paper, we present a simple yet surprisingly powerful approach for unsupervised learning of CNN.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. Boser, J. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1990
Earlier work this paper cites.
Learning invariance from transformation sequences
P. Foldiak · 1991
Earlier work this paper cites.
The” wake-sleep” algorithm for unsupervised neural networks
G. E. Hinton, P. Dayan, B. J. Frey, and R. M. Neal · 1995
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Slow feature analysis:unsupervised learning of invariances
L. Wiskott and T. J. Sejnowski · 2002
Earlier work this paper cites.
Distinctive Image Features from Scale-Invariant Keypoints
D. Lowe · 2004
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
S. Chopra, R. Hadsell, and Y. LeCun · 2005
Earlier work this paper cites.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
Discovering objects and their location in images
J. Sivic, B. C. Russell, A. A. Efros, A. Zisserman, and W. T. Freeman · 2005
Earlier work this paper cites.
Describing visual scenes using transformed dirichlet processes
E. B. Sudderth, A. Torralba, W. T. Freeman, and A. S. Willsky · 2005
Earlier work this paper cites.
Surf: Speeded up robust features
H. Bay, T. Tuytelaars, and L. V. Gool · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
B. C. Russell, A. A. Efros, J. Sivic, W. T. Freeman, and A. Zisserman · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2007
Earlier work this paper cites.
Deep convolutional ranking for multilabel image annotation
Y. Gong, Y. Jia, T. K. Leung, A. Toshev, and S. Ioffe · 2007
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
M. A. Ranzato, F. J. Huang, Y.-L. Boureau, and Y. LeCun · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P. Manzagol · 2008
Earlier work this paper cites.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng · 2009
Earlier work this paper cites.
Deep learning from temporal coherence in video
H. Mobahi, R. Collobert, and J. Weston · 2009
Cited alongside, same era.
Recognizing indoor scenes
A. Quattoni and A.Torralba · 2009
Cited alongside, same era.
The pascal visual object classes (voc) challenge
M. Everingham, L. V. Gool, C. K. Williams, J. Winn, , and A. Zisserman · 2010
Cited alongside, same era.
Unsupervised learning of invariant features using video
D. Stavens and S. Thrun · 2010
Cited alongside, same era.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Cited alongside, same era.
Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis
Q. V. Le, W. Y. Zou, S. Y. Yeung, and A. Y. Ng · 2011
Cited alongside, same era.
Pedestrian detection with unsupervised multi-stage feature learning
P. Sermanet, K. Kavukcuoglu, S. Chintala, and Y. LeCun · 2013
Later among the works it cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Later among the works it cites.
Context as supervisory signal: Discovering objects with predictable context
C. Doersch, A. Gupta, and A. A. Efros · 2014
Later among the works it cites.
Unfolding an indoor origami world
D. F. Fouhey, A. Gupta, and M. Hebert · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Later among the works it cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The shape boltzmann machine: a strong model of object shape
S. M. A. Eslami, N. Heess, and J. Winn · 2012
Cited alongside, same era.
Discriminative decorrelation for clustering and classification
B. Hariharan, J. Malik, and D. Ramanan · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Q. V. Le, M. A. Ranzato, R. Monga, M. Devin, K. Chen, G. S. Corrado, J. Dean, and A. Y. Ng · 2012
Cited alongside, same era.
Hierarchical face parsing via deep learning
P. Luo, X. Wang, and X. Tang · 2012
Cited alongside, same era.
Indoor segmentation and support inference from RGBD images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Cited alongside, same era.
Later among the works it cites.
Discriminatively trained dense surface normal estimation
L. Ladický, B. Zeisl, and M. Pollefeys · 2014
Later among the works it cites.
X. Liang, S. Liu, Y. Wei, L. Liu, L. Lin, and S. Yan · 2014
Later among the works it cites.
Learning fine-grained image similarity with deep ranking
J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu · 2014
Later among the works it cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Closest in time.
Unsupervised learning of spatiotemporally coherent metrics
R. Goroshin, J. Bruna, J. Tompson, D. Eigen, and Y. LeCun · 2015
Closest in time.
High-speed tracking with kernelized correlation filters
J. F. Henriques, R. Caseiro, P. Martins, and J. Batista · 2015
Closest in time.
Deep metric learning using triplet network
E. Hoffer and N. Ailon · 2015
Closest in time.
Matching-cnn meets knn: Quasi-parametric human parsing
S. Liu, X. Liang, L. Liu, X. Shen, J. Yang, C. Xu, X. Cao, and S. Yan · 2015
Closest in time.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Designing deep networks for surface normal estimation
X. Wang, D. F. Fouhey, and A. Gupta · 2015
Closest in time.
Learning descriptors for object recognition and 3d pose estimation
P. Wohlhart and V. Lepetit · 2015
Closest in time.
Bit-scalable deep hashing with regularized similarity learning for image retrieval and person re-identification
R. Zhang, L. Lin, R. Zhang, W. Zuo, and L. Zhang · 2015
Closest in time.