Fetching the paper…
Reading the bibliography…
By extracting spatial and temporal characteristics in one network, the two-stream ConvNets can achieve the state-of-the-art performance in action recognition.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
A duality based approach for realtime tv-l 1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Improving the fisher kernel for large-scale image classification
F. Perronnin, J. Sánchez, and T. Mensink · 2010
Earlier work this paper cites.
Lp-norm multiple kernel learning
M. Kloft, U. Brefeld, S. Sonnenburg, and A. Zien · 2011
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
N. Srivastava and R. R. Salakhutdinov · 2012
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Multi-view super vector for action recognition
Z. Cai, L. Wang, X. Peng, and Y. Qiao · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Exploring inter-feature and inter-class relationships with deep neural networks for video classification
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Grammar as a foreign language
O. Vinyals, Ł. Kaiser, T. Koo, S. Petrov, I. Sutskever, and G. Hinton · 2015
Later among the works it cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Later among the works it cites.
Towards good practices for very deep two-stream convnets
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Later among the works it cites.
Modeling spatial-temporal clues in a hybrid deep learning framework for video classification
Z. Wu, X. Wang, Y.-G. Jiang, H. Ye, and X. Xue · 2015
Later among the works it cites.
Beyond short snippets: Deep networks for video classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Wu, Y.-G. Jiang, J. Wang, J. Pu, and X. Xue · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Action recognition using visual attention
S. Sharma, R. Kiros, and R. Salakhutdinov · 2015
Cited alongside, same era.
Learning deep trajectory descriptor for action recognition in videos using deep neural networks
Y. Shi, W. Zeng, T. Huang, and Y. Wang · 2015
Cited alongside, same era.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Cited alongside, same era.
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Later among the works it cites.
Exploiting image-trained cnn architectures for unconstrained video classification
S. Zha, F. Luisier, W. Andrews, N. Srivastava, and R. Salakhutdinov · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Closest in time.
Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice
X. Peng, L. Wang, X. Wang, and Y. Qiao · 2016
Closest in time.
Sequential deep trajectory descriptor for action recognition with three-stream cnn
Y. Shi, T. Yonghong, Y. Wang, and T. Huang · 2016
Closest in time.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Closest in time.
Action recognition with joint attention on multi-level deep features
J. Wu, G. Wang, W. Yang, and X. Ji · 2016
Closest in time.