Fetching the paper…
Reading the bibliography…
We present an unsupervised representation learning approach using videos without semantic labels.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Discovering objects and their location in images
J. Sivic, B. C. Russell, A. A. Efros, A. Zisserman, and W. T. Freeman · 2005
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
B. C. Russell, W. T. Freeman, A. A. Efros, J. Sivic, and A. Zisserman · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, H. Larochelle, et al · 2007
Earlier work this paper cites.
Self-taught learning: Transfer learning from unlabeled data
R. Raina, A. Battle, H. Lee, B. Packer, and A. Y. Ng · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Deep learning from temporal coherence in video
H. Mobahi, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Q. V. Le · 2012
Earlier work this paper cites.
Unsupervised discovery of mid-level discriminative patches
S. Singh, A. Gupta, and A. A. Efros · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Mid-level visual element discovery as discriminative mode seeking
C. Doersch, A. Gupta, and A. A. Efros · 2013
Earlier work this paper cites.
Learning discriminative part detectors for image classification and cosegmentation
J. Sun and J. Ponce · 2013
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Learning to see by moving
P. Agrawal, J. Carreira, and J. Malik · 2015
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Learning representations for automatic colorization
G. Larsson, M. Maire, and G. Shakhnarovich · 2016
Later among the works it cites.
Unsupervised visual representation learning by graph-based consistent constraints
D. Li, W.-C. Hung, J.-B. Huang, S. Wang, N. Ahuja, and M.-H. Yang · 2016
Later among the works it cites.
Learning image matching by simply watching video
G. Long, L. Kneip, J. M. Alvarez, H. Li, X. Zhang, and Q. Yu · 2016
Later among the works it cites.
Shuffle and learn: Unsupervised learning using temporal order verification
I. Misra, C. L. Zitnick, and M. Hebert · 2016
Later among the works it cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Later among the works it cites.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Learning image representations tied to ego-motion
D. Jayaraman and K. Grauman · 2015
Cited alongside, same era.
Data-dependent initializations of convolutional neural networks
P. Krähenbühl, C. Doersch, J. Donahue, and T. Darrell · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Cited alongside, same era.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
Later among the works it cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krähenbühl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Later among the works it cites.
Pose from action: Unsupervised learning of pose features based on motion
S. Purushwalkam and A. Gupta · 2016
Later among the works it cites.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Later among the works it cites.
Actions ~ transformations
X. Wang, A. Farhadi, and A. Gupta · 2016
Later among the works it cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Later among the works it cites.
Self-supervised video representation learning with odd-one-out networks
B. Fernando, H. Bilen, E. Gavves, and S. Gould · 2017
Closest in time.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2017
Closest in time.
Split-brain autoencoders: Unsupervised learning by cross-channel prediction
R. Zhang, P. Isola, and A. A. Efros · 2017
Closest in time.