Fetching the paper…
Reading the bibliography…
Actions are more than just movements and trajectories: we cook to eat and we hold a cup to drink from it.
Does the chimpanzee have a theory of mind?
D. Premack and G. Woodruff · 1978
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
What, where and who? classifying events by scene and object recognition
L.-J. Li and L. Fei-Fei · 2007
Earlier work this paper cites.
Hierarchical recognition of human activities interacting with objects
M. S. Ryoo and J. Aggarwal · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Klaser, M. Marszalek, and C. Schmid · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
D. Koller and N. Friedman · 2009
Earlier work this paper cites.
Trajectons: Action recognition through the motion analysis of tracked features
P. Matikainen, M. Hebert, and R. Sukthankar · 2009
Earlier work this paper cites.
Deep boltzmann machines
R. Salakhutdinov and G. E. Hinton · 2009
Earlier work this paper cites.
A survey on vision-based human action recognition
R. Poppe · 2010
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
P. Krähenbühl and V. Koltun · 2011
Earlier work this paper cites.
Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis
Q. V. Le, W. Y. Zou, S. Y. Yeung, and A. Y. Ng · 2011
Earlier work this paper cites.
A survey of vision-based methods for action representation, segmentation and recognition
D. Weinland, R. Ronfard, and E. Boyer · 2011
Earlier work this paper cites.
Recognizing complex events using large margin joint low-level event model
H. Izadinia and M. Shah · 2012
Earlier work this paper cites.
Activity forecasting
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Weakly supervised learning of interactions between humans and objects
A. Prest, C. Schmid, and V. Ferrari · 2012
Earlier work this paper cites.
Script data for attribute-based recognition of composite activities
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele · 2012
Earlier work this paper cites.
Action bank: A high-level representation of activity in video
S. Sadanand and J. J. Corso · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. Roshan Zamir, and M. Shah · 2012
Cited alongside, same era.
Learning latent temporal structure for complex event detection
K. Tang, L. Fei-Fei, and D. Koller · 2012
Cited alongside, same era.
Modeling actions through state changes
A. Fathi and J. M. Rehg · 2013
Cited alongside, same era.
Representing videos using mid-level discriminative patches
A. Jain, A. Gupta, M. Rodriguez, and L. S. Davis · 2013
Cited alongside, same era.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Beyond gaussian pyramid: Multi-skip feature stacking for action recognition
Z. Lan, M. Lin, X. Li, A. G. Hauptmann, and B. Raj · 2015
Later among the works it cites.
Fully connected deep structured networks
A. G. Schwing and R. Urtasun · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Later among the works it cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Later among the works it cites.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2013
Cited alongside, same era.
Action recognition by hierarchical sequence summarization
Y. Song, L.-P. Morency, and R. Davis · 2013
Cited alongside, same era.
Active: Activity concept transitions in video event classification
C. Sun and R. Nevatia · 2013
Cited alongside, same era.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Action recognition with stacked fisher vectors
X. Peng, C. Zou, Y. Qiao, and Q. Peng · 2014
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Later among the works it cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Later among the works it cites.
A discriminative cnn video representation for event detection
Z. Xu, Y. Yang, and A. G. Hauptmann · 2015
Later among the works it cites.
Every moment counts: Dense detailed labeling of actions in complex videos
S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and L. Fei-Fei · 2015
Later among the works it cites.
End-to-end learning of action detection from frame glimpses in videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2015
Later among the works it cites.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Later among the works it cites.
Sympathy for the details: Dense trajectories and hybrid classification architectures for action recognition
C. R. de Souza, A. Gaidon, E. Vig, and A. M. López · 2016
Closest in time.
Spot on: Action localization from pointly-supervised proposals
P. Mettes, J. C. van Gemert, and C. G. Snoek · 2016
Closest in time.
Temporal action localization in untrimmed videos via multi-stage cnns
Z. Shou, D. Wang, and S.-F. Chang · 2016
Closest in time.
Learning visual storylines with skipping recurrent neural networks
G. A. Sigurdsson, X. Chen, and A. Gupta · 2016
Closest in time.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Closest in time.
Predicting motivations of actions by leveraging text
C. Vondrick, D. Oktay, H. Pirsiavash, and A. Torralba · 2016
Closest in time.
Anticipating visual representations from unlabeled video
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Closest in time.
Actions ~ transformations
X. Wang, A. Farhadi, and A. Gupta · 2016
Closest in time.
Towards weakly-supervised action localization
P. Weinzaepfel, X. Martin, and C. Schmid · 2016
Closest in time.
Situation recognition: Visual semantic role labeling for image understanding
M. Yatskar, L. Zettlemoyer, and A. Farhadi · 2016
Closest in time.
Instance-level segmentation with deep densely connected mrfs
Z. Zhang, S. Fidler, and R. Urtasun · 2016
Closest in time.