Fetching the paper…
Reading the bibliography…
We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset.
Convolutional networks for images, speech, and time-series
Y. LeCun and Y. Bengio · 1995
Earlier work this paper cites.
Space-time interest points
I. Laptev and T. Lindeberg · 2003
Earlier work this paper cites.
Behavior recognition via sparse spatio-temporal features
P. Dollar, V. Rabaud, G. Cottrell, and S. Belongie · 2005
Earlier work this paper cites.
Detecting irregularities in images and in video
O. Boiman and M. Irani · 2007
Earlier work this paper cites.
A 3-dimensional sift descriptor and its application to action recognition
P. Scovanner, S. Ali, and M. Shah · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Kläser, M. Marszałek, and C. Schmid · 2008
Earlier work this paper cites.
Visualizing data using t-sne
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Egocentric recognition of handled objects: Benchmark and analysis
X. Ren and M. Philipose · 2009
Earlier work this paper cites.
Boundary learning by optimization with topological constraints
V. Jain, B. Bollmann, M. Richardson, D. Berger, M. Helmstaedter, K. Briggman, W. Denk, J. Bowden, J. Mendenhall, W. Abraham, K. Harris, N. Kasthuri, K. Hayworth, R. Schalek, J. Tapia, J. Lichtman, and H. Seung · 2010
Earlier work this paper cites.
Moving vistas: Exploiting motion for describing scenes
N. Shroff, P. K. Turaga, and R. Chellappa · 2010
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Convolutional networks can learn to generate affinity graphs for image segmentation
S. Turaga, J. Murray, V. Jain, F. Roth, M. Helmstaedter, K. Briggman, W. Denk, and S. Seung · 2010
Earlier work this paper cites.
Large displacement optical flow: Descriptor matching in variational motion estimation
T. Brox and J. Malik · 2011
Earlier work this paper cites.
The one shot similarity metric learning for action recognition
O. Kliper-Grossa, T. Hassner, and L. Wolf · 2011
Earlier work this paper cites.
Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis
Q. V. Le, W. Y. Zou, S. Y. Yeung, and A. Y. Ng · 2011
Earlier work this paper cites.
Dynamic scene understanding: The role of orientation features in space and time in scene classification
K. Derpanis, M. Lecce, K. Daniilidis, and R. Wildes · 2012
Earlier work this paper cites.
Motion interchange patterns for action recognition in unconstrained videos
O. Kliper-Gross, Y. Gurovich, T. Hassner, and L. Wolf · 2012
Cited alongside, same era.
The action similarity labeling challenge
O. Kliper-Gross, T. Hassner, and L. Wolf · 2012
Cited alongside, same era.
Activity forecasting
D. B. Kris M. Kitani, Brian D. Ziebart and M. Hebert · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Action bank: A high-level representation of activity in video
S. Sadanand and J. Corso · 2012
Cited alongside, same era.
UCF101: A dataset of 101 human action classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Cited alongside, same era.
Learning human pose estimation features with convolutional networks
A. Jain, J. Tompson, M. Andriluka, G. W. Taylor, and C. Bregler · 2014
Closest in time.
Modeep: A deep learning framework using motion features for human pose estimation
A. Jain, J. Tompson, Y. LeCun, and C. Bregler · 2014
Closest in time.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Closest in time.
THUMOS challenge: Action recognition with a large number of classes, 2014
Y. Jiang, J. Liu, A. Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Closest in time.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2013
Cited alongside, same era.
Spacetime forests with complementary features for dynamic scene recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2013
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2013
Cited alongside, same era.
Evaluating new variants of motion interchange patterns
Y. Hanani, N. Levy, and L. Wolf · 2013
Cited alongside, same era.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Cited alongside, same era.
Dynamic scene classification: Learning motion descriptors with slow features analysis
C. Theriault, N. Thome, and M. Cord · 2013
Cited alongside, same era.
Z. Lan, M. Lin, X. Li, A. G. Hauptmann, and B. Raj · 2014
Closest in time.
Trecvid’14–an overview of the goals, tasks, data, evaluation and metrics
P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, W. Kraaij, A. Smeaton, and G. Quéenot · 2014
Closest in time.
Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice
X. Peng, L. Wang, X. Wang, and Y. Qiao · 2014
Closest in time.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Closest in time.
Large margin dimensionality reduction for action similarity labeling
Q. P. X. Peng, Y. Qiao and Q. Wang · 2014
Closest in time.
Visualizing and understanding convolutional networks
M. Zeiler and R. Fergus · 2014
Closest in time.
Panda: Pose aligned networks for deep attribute modeling
N. Zhang, M. Paluri, M. Ranzato, T. Darrell, and L. Bourdev · 2014
Closest in time.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Closest in time.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.