Fetching the paper…
Reading the bibliography…
Current action recognition methods heavily rely on trimmed videos for model training.
Solving the multiple instance problem with axis-parallel rectangles
T. G. Dietterich, R. H. Lathrop, and T. Lozano-Pérez · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Action snippets: How many frames does human action recognition require?
K. Schindler and L. Van Gool · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li · 2009
Earlier work this paper cites.
Automatic annotation of human actions in video
O. Duchenne, I. Laptev, J. Sivic, F. R. Bach, and J. Ponce · 2009
Earlier work this paper cites.
Modeling the temporal extent of actions
S. Satkin and M. Hebert · 2010
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. A. Poggio, and T. Serre · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Finding actors and actions in movies
P. Bojanowski, F. R. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2013
Earlier work this paper cites.
3D convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Motionlets: Mid-level 3D parts for human motion recognition
L. Wang, Y. Qiao, and X. Tang · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
P. Bojanowski, R. Lajugie, F. R. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2014
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes, 2014
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu · 2014
Cited alongside, same era.
The lear submission at thumos 2014
D. Oneata, J. Verbeek, and C. Schmid · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
CNN: single-label to multi-label
Y. Wei, W. Xia, J. Huang, B. Ni, J. Dong, Y. Zhao, and S. Yan · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Later among the works it cites.
Webly-supervised video recognition by mutually voting for relevant web images and web video frames
C. Gan, C. Sun, L. Duan, and B. Gong · 2016
Later among the works it cites.
You lead, we exceed: Labor-free video concept learning by jointly exploiting web videos and images
C. Gan, T. Yao, K. Yang, Y. Yang, and T. Mei · 2016
Later among the works it cites.
Connectionist temporal modeling for weakly supervised action labeling
D. Huang, L. Fei-Fei, and J. C. Niebles · 2016
Later among the works it cites.
Weakly supervised learning of actions from transcripts
H. Kuehne, A. Richard, and J. Gall · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ActivityNet: A large-scale video benchmark for human activity understanding
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
What do 15, 000 object categories tell us about classifying and localizing actions?
M. Jain, J. C. van Gemert, and C. G. M. Snoek · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
Is object localization for free? - weakly-supervised learning with convolutional neural networks
M. Oquab, L. Bottou, I. Laptev, and J. Sivic · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li · 2015
Cited alongside, same era.
Learning activity progression in LSTMs for activity detection and early detection
S. Ma, L. Sigal, and S. Sclaroff · 2016
Later among the works it cites.
Temporal action detection using a statistical language model
A. Richard and J. Gall · 2016
Later among the works it cites.
Temporal action localization in untrimmed videos via multi-stage CNNs
Z. Shou, D. Wang, and S. Chang · 2016
Later among the works it cites.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2016
Later among the works it cites.
MoFAP: A multi-level representation for action recognition
L. Wang, Y. Qiao, and X. Tang · 2016
Later among the works it cites.
Actionness estimation using hybrid fully convolutional networks
L. Wang, Y. Qiao, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Val Gool · 2016
Later among the works it cites.
CUHK & ETHZ & SIAT submission to ActivityNet challenge 2016
Y. Xiong, L. Wang, Z. Wang, B. Zhang, H. Song, W. Li, D. Lin, Y. Qiao, L. Van Gool, and X. Tang · 2016
Later among the works it cites.
End-to-end learning of action detection from frame glimpses in videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2016
Later among the works it cites.
Temporal action localization with pyramid of score distribution features
J. Yuan, B. Ni, X. Yang, and A. A. Kassim · 2016
Later among the works it cites.
Real-time action recognition with enhanced motion vector CNNs
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang · 2016
Later among the works it cites.
Depth2action: Exploring embedded depth for large-scale action recognition
Y. Zhu and S. D. Newsam · 2016
Later among the works it cites.
Weakly supervised object localization with multi-fold multiple instance learning
R. G. Cinbis, J. J. Verbeek, and C. Schmid · 2017
Closest in time.