Fetching the paper…
Reading the bibliography…
In this paper, we present an approach for learning a visual representation from the raw spatiotemporal signals in videos.
Communication in the presence of noise
Shannon, C.E.: · 1949
Earlier work this paper cites.
A synopsis of linguistic theory 1930-1955
Firth, J.R.: · 1957
Earlier work this paper cites.
Implicit learning and tacit knowledge
Reber, A.S.: · 1989
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J.S., Henderson, D., Howard, R.E., Hubbard, W., Jackel, L.D.: · 1989
Earlier work this paper cites.
Learning the structure of event sequences
Cleeremans, A., McClelland, J.L.: · 1991
Earlier work this paper cites.
Learning invariance from transformation sequences
Földiák, P.: · 1991
Earlier work this paper cites.
Mechanisms of implicit learning: Connectionist models of sequence processing
Cleeremans, A.: · 1993
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Olshausen, B.A., et al.: · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J.: · 1997
Earlier work this paper cites.
From implicit skills to explicit knowledge: A bottom-up model of skill learning
Sun, R., Merrill, E., Peterson, T.: · 2001
Earlier work this paper cites.
Sequence learning: from recognition and prediction to sequential decision making
Sun, R., Giles, C.L.: · 2001
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Wiskott, L., Sejnowski, T.J.: · 2002
Earlier work this paper cites.
Two-frame motion estimation based on polynomial expansion
Farnebäck, G.: · 2003
Earlier work this paper cites.
Discovering objects and their location in images
Sivic, J., Russell, B.C., Efros, A.A., Zisserman, A., Freeman, W.T.: · 2005
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
Russell, B.C., Freeman, W.T., Efros, A.A., Sivic, J., Zisserman, A.: · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Lee, H., Battle, A., Raina, R., Ng, A.Y.: · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R., Chopra, S., LeCun, Y.: · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., Larochelle, H., et al.: · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., Manzagol, P.A.: · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., jia Li, L., Li, K., Fei-fei, L.: · 2009
Earlier work this paper cites.
Deep boltzmann machines
Salakhutdinov, R., Hinton, G.E.: · 2009
Earlier work this paper cites.
Deep learning from temporal coherence in video
Mobahi, H., Collobert, R., Weston, J.: · 2009
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
Taylor, G.W., Fergus, R., LeCun, Y., Bregler, C.: · 2010
Earlier work this paper cites.
A survey on vision-based human action recognition
Poppe, R.: · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., Singer, Y.: · 2011
Cited alongside, same era.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M.: · 2012
Cited alongside, same era.
“clustering by composition”–unsupervised discovery of image categories
Faktor, A., Irani, M.: · 2012
Cited alongside, same era.
Unsupervised discovery of mid-level discriminative patches
Singh, S., Gupta, A., Efros, A.: · 2012
Cited alongside, same era.
Slow feature analysis for human action recognition
Zhang, Z., Tao, D.: · 2012
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Later among the works it cites.
Return of the devil in the details: Delving deep into convolutional nets
Chatfield, K., Simonyan, K., Vedaldi, A., Zisserman, A.: · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., Malik, J.: · 2014
Later among the works it cites.
Deeppose: Human pose estimation via deep neural networks
Toshev, A., Szegedy, C.: · 2014
Later among the works it cites.
Unsupervised visual representation learning by context prediction
Doersch, C., Gupta, A., Efros, A.A.: · 2015
Later among the works it cites.
Learning image representations equivariant to ego-motion
Jayaraman, D., Grauman, K.: · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E.: · 2012
Cited alongside, same era.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., Dean, J.: · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: · 2013
Cited alongside, same era.
Modec: Multimodal decomposable models for human pose estimation
Sapp, B., Taskar, B.: · 2013
Cited alongside, same era.
Blocks that shout: Distinctive parts for scene classification
Juneja, M., Vedaldi, A., Jawahar, C., Zisserman, A.: · 2013
Cited alongside, same era.
Mid-level visual element discovery as discriminative mode seeking
Doersch, C., Gupta, A., Efros, A.A.: · 2013
Cited alongside, same era.
Later among the works it cites.
Slow and steady feature analysis: Higher order temporal coherence in video
Jayaraman, D., Grauman, K.: · 2015
Later among the works it cites.
Learning visual groups from co-occurrences in space and time
Isola, P., Zoran, D., Krishnan, D., Adelson, E.H.: · 2015
Later among the works it cites.
Unsupervised learning of spatiotemporally coherent metrics
Goroshin, R., Bruna, J., Tompson, J., Eigen, D., LeCun, Y.: · 2015
Later among the works it cites.
Unsupervised learning of video representations using lstms
Srivastava, N., Mansimov, E., Salakhutdinov, R.: · 2015
Later among the works it cites.
Temporal perception and prediction in ego-centric video
Zhou, Y., Berg, T.L.: · 2015
Later among the works it cites.
Anticipating the future by watching unlabeled video
Vondrick, C., Pirsiavash, H., Torralba, A.: · 2015
Later among the works it cites.
Learning to see by moving
Agrawal, P., Carreira, J., Malik, J.: · 2015
Later among the works it cites.
Unsupervised learning of visual representations using videos
Wang, X., Gupta, A.: · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., Szegedy, C.: · 2015
Later among the works it cites.
Towards good practices for very deep two-stream convnets
Wang, L., Xiong, Y., Wang, Z., Qiao, Y.: · 2015
Later among the works it cites.
Fast R-CNN
Girshick, R.: · 2015
Later among the works it cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., Sun, J.: · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: · 2015
Later among the works it cites.
Watch and learn: Semi-supervised learning of object detectors from videos
Misra, I., Shrivastava, A., Hebert, M.: · 2015
Later among the works it cites.
Towards computational baby learning: A weakly-supervised approach for object detection
Liang, X., Liu, S., Wei, Y., Liu, L., Lin, L., Yan, S.: · 2015
Later among the works it cites.
Flowing convnets for human pose estimation in videos
Pfister, T., Charles, J., Zisserman, A.: · 2015
Later among the works it cites.
Generative image modeling using style and structure adversarial networks
Wang, X., Gupta, A.: · 2016
Closest in time.
Visually indicated sounds
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E.H., Freeman, W.T.: · 2016
Closest in time.