Fetching the paper…
Reading the bibliography…
Every moment counts in action recognition.
Recognizing human action in time-sequential images using hidden markov model
J. Yamato, J. Ohya, and K. Ishii · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
A model of saliency-based visual attention for rapid scene analysis
L. Itti, C. Koch, and E. Niebur · 1998
Earlier work this paper cites.
Recognizing human actions: A local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Actions as space-time shapes
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri · 2005
Earlier work this paper cites.
Event detection in crowded videos
Y. Ke, R. Sukthankar, and M. Hebert · 2007
Earlier work this paper cites.
Single view human action recognition using key pose matching and viterbi path searching
F. J. Lv and R. Nevatia · 2007
Earlier work this paper cites.
Actions in context
M. Marszałek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Spatio-temporal relationship match: Video structure comparison for recognition of complex human activities
M. S. Ryoo and J. K. Aggarwal · 2009
Earlier work this paper cites.
Exploiting hierarchical context on a large database of object categories
M. J. Choi, J. J. Lim, A. Torralba, and A. S. Willsky · 2010
Earlier work this paper cites.
Modeling temporal structure of decomposable motion segments for activity classification
J. C. Niebles, C.-W. Chen, and L. Fei-Fei · 2010
Earlier work this paper cites.
A survey on vision-based human action recognition
R. Poppe · 2010
Earlier work this paper cites.
A survey of vision-based methods for action representation, segmentation and recognition
D. Weinland, R. Ronfard, and E. Boyer · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Discriminative figure-centric models for joint action localization and recognition
T. Lan, Y. Wang, and G. Mori · 2011
Earlier work this paper cites.
A large-scale benchmark dataset for event recognition in surveillance video
S. Oh, A. Hoogs, A. Perera, N. Cuntoor, C.-C. Chen, J. T. Lee, S. Mukherjee, J. K. Aggarwal, H. Lee, L. Davis, E. Swears, X. Wang, Q. Ji, K. Reddy, M. Shah, C. Vondrick, H. Pirsiavash, D. Ramanan, J. Yuen, A. Torralba, B. Song, A. Fong, A. Roy-Chowdhury, and M. Desai · 2011
Earlier work this paper cites.
Trecvid 2011 — an overview of the goals, tasks, data, evaluation mechansims and metrics
P. Over, G. Awad, M. Michel, J. Fiscus, W. Kraaij, A. F. Smeaton, and G. Quenot · 2011
Earlier work this paper cites.
Human action segmentation and recognition using discriminative semi-markov models
L. W. A. S. Qinfeng Shi, Li Cheng · 2011
Earlier work this paper cites.
A discriminative key pose sequence model for recognizing human interactions
A. Vahdat, B. Gao, M. Ranjbar, and G. Mori · 2011
Cited alongside, same era.
Action recognition by dense trajectories
H. Wang, A. Kläser, C.Schmid, and C.-L. Liu · 2011
Cited alongside, same era.
Activity forecasting
K. M. Kitani, B. Ziebart, J. D. Bagnell, and M. Hebert · 2012
Cited alongside, same era.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Cited alongside, same era.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Cited alongside, same era.
Learning latent temporal structure for complex event detection
K. Tang, L. Fei-Fei, and D. Koller · 2012
Multiple granularity analysis for fine-grained action detection
B. Ni, V. R. Paramathayalan, and P. Moulin · 2014
Later among the works it cites.
Multimedia event detection with multimodal feature fusion and temporal concept localization
S. Oh, S. Mccloskey, I. Kim, A. Vahdat, K. J. Cannons, H. Hajimirsadeghi, G. Mori, A. A. Perera, M. Pandey, and J. J. Corso · 2014
Later among the works it cites.
Parsing videos of actions with segmental grammars
H. Pirsiavash and D. Ramanan · 2014
Later among the works it cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude., 2012
T. Tieleman and G. E. Hinton · 2012
Cited alongside, same era.
Action is in the eye of the beholder: Eye-gaze driven model for spatio-temporal action localization
N. Shapovalova, M. Raptis, L. Sigal, and G. Mori · 2013
Cited alongside, same era.
Spatiotemporal deformable part models for action detection
Y. Tian, R. Sukthankar, and M. Shah · 2013
Cited alongside, same era.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
W. Tong, Y. Yang, L. Jiang, S.-I. Yu, Z. Lan, Z. Ma, W. Sze, E. Younessian, and A. G. Hauptmann · 2014
Later among the works it cites.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Closest in time.
Initialization strategies of spatio-temporal convolutional neural networks
E. Mansimov, N. Srivastava, and R. Salakhutdinov · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.
Recognizing fine-grained and composite activities using hand-centric features and script data
M. Rohrbach, A. Rohrbach, M. Regneri, S. Amin, M. Andriluka, M. Pinkal, and B. Schiele · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky et al · 2015
Closest in time.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
C3d: Generic features for video analysis
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Closest in time.
Temporal pyramid pooling based convolutional neural networks for action recognition
P. Wang, Y. Cao, C. Shen, L. Liu, and H. T. Shen · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu et al · 2015
Closest in time.
Video description generation incorporating spatio-temporal features and a soft-attention mechanism
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.
Exploiting image-trained cnn architectures for unconstrained video classification
S. Zha, F. Luisier, W. Andrews, N. Srivastava, and R. Salakhutdinov · 2015
Closest in time.