Fetching the paper…
Reading the bibliography…
Supervised (pre-)training currently yields state-of-the-art performance for representation learning for visual recognition, yet it comes at the cost of (1) intensive manual annotations and (2) an inherent restriction in the scope of data relevant for learning.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Olshausen, B.A., Field, D.J.: · 1997
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Wiskott, L., Sejnowski, T.J.: · 2002
Earlier work this paper cites.
Simple-cell-like receptive fields maximize temporal coherence in natural video
Hurri, J., Hyvärinen, A.: · 2003
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G.E., Salakhutdinov, R.R.: · 2006
Earlier work this paper cites.
Surf: Speeded up robust features
Bay, H., Tuytelaars, T., Van Gool, L.: · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., Larochelle, H., et al.: · 2007
Earlier work this paper cites.
Learning globally-consistent local distance functions for shape-based image retrieval and classification
Frome, A., Singer, Y., Sha, F., Malik, J.: · 2007
Earlier work this paper cites.
Deep learning from temporal coherence in video
Mobahi, H., Collobert, R., Weston, J.: · 2009
Earlier work this paper cites.
Learning deep architectures for ai
Bengio, Y.: · 2009
Earlier work this paper cites.
Slow, decorrelated features for pretraining complex cell-like networks
Bengio, Y., Bergstra, J.S.: · 2009
Earlier work this paper cites.
Quattoni, A., Torralba, A.: · 2009
Earlier work this paper cites.
Unsupervised learning of visual invariance with temporal coherence
Zou, W.Y., Ng, A.Y., Yu, K.: · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E.: · 2012
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Le, Q.V.: · 2012
Cited alongside, same era.
Deep learning of invariant features via simulated fixations in video
Zou, W., Zhu, S., Yu, K., Ng, A.Y.: · 2012
Cited alongside, same era.
Selective search for object recognition
Uijlings, J.R., van de Sande, K.E., Gevers, T., Smeulders, A.W.: · 2013
Cited alongside, same era.
Action recognition with improved trajectories
Wang, H., Schmid, C.: · 2013
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., Darrell, T.: · 2014
Cited alongside, same era.
Unsupervised feature learning from temporal data
Goroshin, R., Bruna, J., Tompson, J., Eigen, D., LeCun, Y.: · 2015
Later among the works it cites.
Unsupervised visual representation learning by context prediction
Doersch, C., Gupta, A., Efros, A.A.: · 2015
Later among the works it cites.
Learning to see by moving
Agrawal, P., Carreira, J., Malik, J.: · 2015
Later among the works it cites.
Learning image representations equivariant to ego-motion
Jayaraman, D., Grauman, K.: · 2015
Later among the works it cites.
Unsupervised learning of video representations using lstms
Srivastava, N., Mansimov, E., Salakhutdinov, R.: · 2015
Later among the works it cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., Philbin, J.: · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., Malik, J.: · 2014
Cited alongside, same era.
Learning fine-grained image similarity with deep ranking
Wang, J., Song, Y., Leung, T., Rosenberg, C., Wang, J., Philbin, J., Chen, B., Wu, Y.: · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., Darrell, T.: · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
Wang, X., Gupta, A.: · 2015
Cited alongside, same era.
Learning temporal embeddings for complex video analysis
Ramanathan, V., Tang, K., Mori, G., Fei-Fei, L.: · 2015
Cited alongside, same era.
Towards computational baby learning: A weakly-supervised approach for object detection
Liang, X., Liu, S., Wei, Y., Liu, L., Lin, L., Yan, S.: · 2015
Later among the works it cites.
High-speed tracking with kernelized correlation filters
Henriques, J.F., Caseiro, R., Martins, P., Batista, J.: · 2015
Later among the works it cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Closest in time.
Slow and steady feature analysis: Higher order temporal coherence in video
Jayaraman, D., Grauman, K.: · 2016
Closest in time.
Context encoders: Feature learning by inpainting
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., Efros, A.A.: · 2016
Closest in time.
Deep metric learning via lifted structured feature embedding
Song, H.O., Xiang, Y., Jegelka, S., Savarese, S.: · 2016
Closest in time.