Fetching the paper…
Reading the bibliography…
In recent years, deep neural network approaches have naturally extended to the video domain, in their simplest case by aggregating per-frame classifications as a baseline for action recognition.
Two-frame motion estimation based on polynomial expansion
G. Farnebäck · 2003
Earlier work this paper cites.
Action mach a spatio-temporal maximum average correlation height filter for action recognition
M. D. Rodriguez, J. Ahmed, and M. Shah · 2008
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Articulated people detection and pose estimation: Reshaping the future
L. Pishchulin, A. Jain, M. Andriluka, T. Thormählen, and B. Schiele · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Deep learning of invariant features via simulated fixations in video
W. Zou, S. Zhu, K. Yu, and A. Y. Ng · 2012
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Using k-poselets for detecting people and localizing their keypoints
G. Gkioxari, B. Hariharan, R. Girshick, and J. Malik · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Deeppose: Human pose estimation via deep neural networks
A. Toshev and C. Szegedy · 2014
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2014
Cited alongside, same era.
3d human activity recognition with reconfigurable convolutional neural networks
K. Wang, X. Wang, L. Lin, M. Wang, and W. Zuo · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi and A. A. et. al · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Later among the works it cites.
Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Later among the works it cites.
Towards good practices for very deep two-stream convnets
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Later among the works it cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Chéron, I. Laptev, and C. Schmid · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Spatio-temporal lstm with trust gates for 3d human action recognition
J. Liu, A. Shahroudy, D. Xu, and G. Wang · 2016
Later among the works it cites.
S.-E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh · 2016
Later among the works it cites.
Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh · 2017
Later among the works it cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
C. Pantofaru, C. Sun, C. Gu, C. Schmid, D. Ross, G. Toderici, J. Malik, R. Sukthankar, S. Vijayanarasimhan, S. Ricco, et al · 2017
Later among the works it cites.