Fetching the paper…
Reading the bibliography…
What defines an action like "kicking ball"? We argue that the true meaning of an action lies in the change or transformation an action brings to the environment.
Signature verification using a “siamese” time delay neural network
J. Bromley, I. Guyon, Y. LeCun, E. Sackinger, and R. Shah · 1993
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
S. Chopra, R. Hadsell, and Y. LeCun · 2005
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
A duality based approach for realtime tv-l1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Klaser, M. Marszalek, and C. Schmid · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Trajectons: Action recognition through the motion analysis of tracked features
P. Matikainen, M. Hebert, and R. Sukthankar · 2009
Earlier work this paper cites.
A survey on vision-based human action recognition
R. Poppe · 2010
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Consumer video understanding: A benchmark database and an evaluation of human and machine performance
Y.-G. Jiang, G. Ye, S.-F. Chang, D. Ellis, and A. C. Loui · 2011
Earlier work this paper cites.
Hmdb: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T.Serre · 2011
Earlier work this paper cites.
Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis
Q. V. Le, W. Y. Zou, S. Y. Yeung, and A. Y. Ng · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
H. Wang, A. Klaser, C. Schmid, and L. Cheng-Lin · 2011
Earlier work this paper cites.
Hidden part models for human action recognition: Probabilistic vs. max-margin
Y. Wang and G. Mori · 2011
Earlier work this paper cites.
Recognizing complex events using large margin joint low-level event model
H. Izadinia and M. Shah · 2012
Earlier work this paper cites.
Trajectory-based modeling of human actions with motion reference points
Y.-G. Jiang, Q. Dai, X. Xue, W. Liu, and C.-W. Ngo · 2012
Earlier work this paper cites.
Script data for attribute-based recognition of composite activities
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele · 2012
Earlier work this paper cites.
Action bank: A high-level representation of activity in video
S. Sadanand and J. J. Corso · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Cited alongside, same era.
Learning latent temporal structure for complex event detection
K. Tang, L. Fei-Fei, and D. Koller · 2012
Cited alongside, same era.
Modeling actions through state changes
A. Fathi and J. M. Rehg · 2013
Cited alongside, same era.
Representing videos using mid-level discriminative patches
A. Jain, A. Gupta, M. Rodriguez, and L. S. Davis · 2013
Cited alongside, same era.
Better exploiting motion for better action recognition
M. Jain, H. Jegou, and P. Bouthemy · 2013
Cited alongside, same era.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, J. O. M., A. Ghodrati, and T. Tuytelaars · 2015
Closest in time.
Devnet: A deep event network for multimedia event detection and evidence recounting
C. Gan, N. Wang, Y. Yang, D.-Y. Yeung, and A. G. Hauptmann · 2015
Closest in time.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Closest in time.
Activitynet: A large-scale video benchmark for human activity understanding
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Closest in time.
Learning image representations tied to ego-motion
D. Jayaraman and K. Grauman · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2013
Cited alongside, same era.
Action recognition by hierarchical sequence summarization
Y. Song, L.-P. Morency, and R. Davis · 2013
Cited alongside, same era.
Active: Activity concept transitions in video event classification
C. Sun and R. Nevatia · 2013
Cited alongside, same era.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Action recognition with actons
J. Zhu, B. Wang, X. Yang, W. Zhang, and Z. Tu · 2013
Cited alongside, same era.
Deep metric learning using triplet network
E. Hoffer and N. Ailon · 2014
Cited alongside, same era.
Action recognition by hierarchical mid-level action elements
T. Lan, Y. Zhu, A. R. Zamir, and S. Savarese · 2015
Closest in time.
Beyond gaussian pyramid: Multi-skip feature stacking for action recognition
Z. Lan, M. Lin, X. Li, A. G. Hauptmann, and B. Raj · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Closest in time.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Temporal localization of fine-grained actions in videos by domain transfer from web images
C. Sun, S. Shetty, R. Sukthankar, and R. Nevatia · 2015
Closest in time.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Closest in time.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Closest in time.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Closest in time.
Towards good practices for very deep two-stream convnets
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Closest in time.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Closest in time.
Modeling spatial-temporal clues in a hybrid deep learning framework for video classification
Z. Wu, X. Wang, Y.-G. Jiang, H. Ye, and X. Xue · 2015
Closest in time.
A discriminative cnn video representation for event detection
Z. Xu, Y. Yang, and A. G. Hauptmann · 2015
Closest in time.