Fetching the paper…
Reading the bibliography…
Recently, attempts have been made to collect millions of videos to train CNN models for action recognition in videos.
Learning realistic human actions from movies
I. Laptev, M. Marszałek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Learning actions from the web
N. Ikizler-Cinbis, R. G. Cinbis, and S. Sclaroff · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Improving the fisher kernel for large-scale image classification
F. Perronnin, J. Sánchez, and T. Mensink · 2010
Earlier work this paper cites.
Grouplet: A structured image representation for recognizing human and object interactions
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Hidden part models for human action recognition: Probabilistic versus max margin
Y. Wang and G. Mori · 2011
Earlier work this paper cites.
A survey of vision-based methods for action representation, segmentation and recognition
D. Weinland, R. Ronfard, and E. Boyer · 2011
Earlier work this paper cites.
Human action recognition by learning bases of action attributes and parts
B. Yao, X. Jiang, A. Khosla, A. L. Lin, L. Guibas, and L. Fei-Fei · 2011
Earlier work this paper cites.
Exploiting web images for event recognition in consumer videos: A multiple source domain adaptation approach
L. Duan, D. Xu, and S.-F. Chang · 2012
Earlier work this paper cites.
Web-based classifiers for human action recognition
N. Ikizler-Cinbis and S. Sclaroff · 2012
Cited alongside, same era.
UCF101: A dataset of 101 human action classes from videos in the wild
A. R. Z. Khurram Soomro and M. Shah · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Cited alongside, same era.
Watching unlabeled video helps learn new human actions from very few labeled snapshots
C.-Y. Chen and K. Grauman · 2013
Cited alongside, same era.
3d convolutional neural networks for human action recognition
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Later among the works it cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Later among the works it cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Video annotation via image groups from the web
H. Wang, X. Wu, and Y. Jia · 2014
Later among the works it cites.
Video action detection with relational dynamic-poselets
L. Wang, Y. Qiao, and X. Tang · 2014
Later among the works it cites.
Activitynet: A large-scale video benchmark for human activity understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Cited alongside, same era.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, I. Laptev, M. Piccardi, M. Shah, and R. Sukthankar · 2013
Cited alongside, same era.
Poselet key-framing: A model for human activity recognition
M. Raptis and L. Sigal · 2013
Cited alongside, same era.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Lear-inria submission for the thumos workshop
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Return of the devil in the details: Delving deep into convolutional nets
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Cited alongside, same era.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Closest in time.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Closest in time.
Temporal localization of fine-grained actions in videos by domain transfer from web images
C. Sun, S. Shetty, R. Sukthankar, and R. Nevatia · 2015
Closest in time.