Fetching the paper…
Reading the bibliography…
Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation.
The application of hidden markov models in speech recognition
Mark Gales and Steve Young · 1932
Earlier work this paper cites.
Modeling by shortest data description
J. Rissanen · 1978
Earlier work this paper cites.
Representation and recognition of the movements of shapes
D. Marr and Lucia Vaina · 1982
Earlier work this paper cites.
Model-based vision: a program to see a walking person
David Hogg · 1983
Earlier work this paper cites.
A combined corner and edge detector
Chris Harris and Mike Stephens · 1988
Earlier work this paper cites.
Static and dynamic error propagation networks with application to speech coding
A. J. Robinson and F. Fallside · 1988
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
L. R. Rabiner · 1989
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Towards model-based recognition of human movements in image sequences
K. Rohr · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A. Olshausen and David J. Field · 1997
Earlier work this paper cites.
Exploiting generative models in discriminative classifiers
Tommi Jaakkola and David Haussler · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
The recognition of human movement using temporal templates
A. F. Bobick and J. W. Davis · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira · 2001
Earlier work this paper cites.
Multiresolution gray-scale and rotation invariant texture classification with local binary patterns
T. Ojala, M. Pietikainen, and T. Maenpaa · 2002
Earlier work this paper cites.
Dynamic textures
Gianfranco Doretto, Alessandro Chiuso, Ying Nian Wu, and Stefano Soatto · 2003
Earlier work this paper cites.
1 2 Separate Visual Pathways for Perception and Action
Melvyn A. Goodale and A. David Milner · 2003
Earlier work this paper cites.
Large-scale event detection using semi-hidden markov models
Somboon Hongeng and Ramakant Nevatia · 2003
Earlier work this paper cites.
Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons
Andrew McCallum and Wei Li · 2003
Earlier work this paper cites.
Visual categorization with bags of keypoints
Gabriella Csurka, Christopher R. Dance, Lixin Fan, Jutta Willamowski, and Cédric Bray · 2004
Earlier work this paper cites.
Recognizing human actions: A local svm approach
Christian Schuldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
Actions as space-time shapes
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri · 2005
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, and Yann LeCun · 2005
Earlier work this paper cites.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
Behavior recognition via sparse spatio-temporal features
P. Dollar, V. Rabaud, G. Cottrell, and S. Belongie · 2005
Earlier work this paper cites.
A bayesian hierarchical model for learning natural scene categories
L. Fei-Fei and P. Perona · 2005
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
Actions sketch: a novel action representation
Alper Yilmaz and Mubarak Shah · 2005
Earlier work this paper cites.
Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information
Emmanuel J. Candès, Justin Romberg, and Terence Tao · 2006
Earlier work this paper cites.
Human Detection Using Oriented Histograms of Flow and Appearance
Navneet Dalal, Bill Triggs, and Cordelia Schmid · 2006
Earlier work this paper cites.
Compressed sensing
David L. Donoho · 2006
Earlier work this paper cites.
A survey of advances in vision-based human motion capture and analysis
Thomas B. Moeslund and Erik Granum · 2006
Earlier work this paper cites.
Sampling strategies for bag-of-features image classification
Eric Nowak, Frédéric Jurie, and Bill Triggs · 2006
Earlier work this paper cites.
Region covariance: A fast descriptor for detection and classification
Oncel Tuzel, Fatih Porikli, and Peter Meer · 2006
Earlier work this paper cites.
Free viewpoint action recognition using motion history volumes
Daniel Weinland, Remi Ronfard, and Edmond Boyer · 2006
Earlier work this paper cites.
Object tracking: A survey
Alper Yilmaz, Omar Javed, and Mubarak Shah · 2006
Earlier work this paper cites.
Fisher kernels on visual vocabularies for image categorization
F. Perronnin and C. Dance · 2007
Earlier work this paper cites.
Hidden conditional random fields
Ariadna Quattoni, Sybor Wang, Louis-Philippe Morency, Michael Collins, and Trevor Darrell · 2007
Earlier work this paper cites.
Dynamic texture recognition using local binary patterns with an application to facial expressions
G. Zhao and M. Pietikainen · 2007
Earlier work this paper cites.
Human activity recognition using a dynamic texture based method
Vili Kellokumpu, Guoying Zhao, and Matti Pietikäinen · 2008
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Kläser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Sparse representation for color image restoration
Julien Mairal, Michael Elad, and Guillermo Sapiro · 2008
Earlier work this paper cites.
Action recognition with motion-appearance vocabulary forest
K. Mikolajczyk and H. Uemura · 2008
Earlier work this paper cites.
Action mach a spatio-temporal maximum average correlation height filter for action recognition
M. D. Rodriguez, J. Ahmed, and M. Shah · 2008
Earlier work this paper cites.
Machine recognition of human activities: A survey
P. Turaga, R. Chellappa, V. S. Subrahmanian, and O. Udrea · 2008
Earlier work this paper cites.
Pedestrian detection via classification on riemannian manifolds
O. Tuzel, F. Porikli, and P. Meer · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
An efficient dense and scale-invariant spatio-temporal interest point detector
Geert Willems, Tinne Tuytelaars, and Luc Gool · 2008
Earlier work this paper cites.
Crowd analysis: a survey
Beibei Zhan, Dorothy N. Monekosso, Paolo Remagnino, Sergio A. Velastin, and Li-Qun Xu · 2008
Earlier work this paper cites.
Recognizing realistic actions from videos ”in the wild”
J. Liu, Jiebo Luo, and M. Shah · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Cited alongside, same era.
Trajectons: Action recognition through the motion analysis of tracked features
P. Matikainen, M. Hebert, and R. Sukthankar · 2009
Cited alongside, same era.
Activity recognition using the velocity histories of tracked keypoints
R. Messing, C. Pal, and H. Kautz · 2009
Cited alongside, same era.
Evaluation of local spatio-temporal features for action recognition
Heng Wang, Muhammad Muneeb Ullah, Alexander Kläser, Ivan Laptev, and Cordelia Schmid · 2009
Cited alongside, same era.
Robust face recognition via sparse representation
John Wright, Allen Y. Yang, Arvind Ganesh, Shankar S. Sastry, and Yi Ma · 2009
Cited alongside, same era.
Sparse and Redundant Representations - From Theory to Applications in Signal and Image Processing
Michael Elad · 2010
Asymmetric and category invariant feature transformations for domain adaptation
Judy Hoffman, Erik Rodner, Jeff Donahue, Brian Kulis, and Kate Saenko · 2014
Later among the works it cites.
Efficient feature extraction, encoding, and classification for action recognition
V. Kantorov and I. Laptev · 2014
Later among the works it cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Later among the works it cites.
Video (language) modeling: a baseline for generative models of natural videos
MarcAurelio Ranzato, Arthur Szlam, Joan Bruna, Michael Mathieu, Ronan Collobert, and Sumit Chopra · 2014
Later among the works it cites.
Spatio-temporal laplacian pyramid coding for action recognition
L. Shao, X. Zhen, D. Tao, and X. Li · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Aggregating local descriptors into a compact image representation
H. Jégou, M. Douze, C. Schmid, and P. Pérez · 2010
Cited alongside, same era.
Learning a hierarchy of discriminative space-time neighborhood features for human action recognition
A. Kovashka and K. Grauman · 2010
Cited alongside, same era.
Object bank: A high-level image representation for scene classification & semantic feature sparsification
Li-Jia Li, Hao Su, Li Fei-fei, and Eric P. Xing · 2010
Cited alongside, same era.
Modeling temporal structure of decomposable motion segments for activity classification
Juan Carlos Niebles, Chih-Wei Chen, and Li Fei-Fei · 2010
Cited alongside, same era.
A survey on vision-based human action recognition
Ronald Poppe · 2010
Cited alongside, same era.
Distributed Video Sensor Networks , chapter Motion Analysis: Past, Present and Future, pages 27–39
J. K. Aggarwal · 2011
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Later among the works it cites.
Action recognition using global spatio-temporal features derived from sparse representations
Guruprasad Somasundaram, Anoop Cherian, Vassilios Morellas, and Nikolaos Papanikolopoulos · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Later among the works it cites.
Human action recognition by representing 3d skeletons as points in a lie group
R. Vemulapalli, F. Arrate, and R. Chellappa · 2014
Later among the works it cites.
Towards good practices for action video encoding
J. Wu, Y. Zhang, and W. Lin · 2014
Later among the works it cites.
Modeling video dynamics with deep dynencoder
Xing Yan, Hong Chang, Shiguang Shan, and Xilin Chen · 2014
Later among the works it cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Later among the works it cites.
Free-form region description with second-order pooling
J. Carreira, R. Caseiro, J. Batista, and C. Sminchisescu · 2015
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Later among the works it cites.
Hierarchical recurrent neural network for skeleton based action recognition
Yong Du, W. Wang, and L. Wang · 2015
Later among the works it cites.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, M. José Oramas, A. Ghodrati, and T. Tuytelaars · 2015
Later among the works it cites.
Unsupervised learning of spatiotemporally coherent metrics
R. Goroshin, J. Bruna, J. Tompson, D. Eigen, and Y. LeCun · 2015
Later among the works it cites.
Improving human action recognition using score distribution and ranking
Minh Hoai and Andrew Zisserman · 2015
Later among the works it cites.
Human action recognition in unconstrained videos by explicit motion modeling
Y. G. Jiang, Q. Dai, W. Liu, X. Xue, and C. W. Ngo · 2015
Later among the works it cites.
Beyond gaussian pyramid: Multi-skip feature stacking for action recognition
Zhenzhong Lan, Ming Lin, Xuanchong Li, A. G. Hauptmann, and B. Raj · 2015
Later among the works it cites.
Deep multi-scale video prediction beyond mean square error
Michaël Mathieu, Camille Couprie, and Yann LeCun · 2015
Later among the works it cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Later among the works it cites.
Learning a non-linear knowledge transfer model for cross-view action recognition
H. Rahmani and A. Mian · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Later among the works it cites.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D. Y. Yeung, and B. E. Shi · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, Wei Liu, Yangqing Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Later among the works it cites.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Later among the works it cites.
Fusing multi-stream deep networks for video classification
Zuxuan Wu, Yu-Gang Jiang, Xi Wang, Hao Ye, Xiangyang Xue, and Jun Wang · 2015
Later among the works it cites.
Action recognition using hybrid feature descriptor and vlad video encoding
Dong Xing, Xianzhong Wang, and Hongtao Lu · 2015
Later among the works it cites.
Johanna Carvajal, Chris McCool, Brian C. Lovell, and Conrad Sanderson · 2016
Closest in time.
Sympathy for the details: Dense trajectories and hybrid classification architectures for action recognition
César Roberto de Souza, Adrien Gaidon, Eleonora Vig, and Antonio Manuel López · 2016
Closest in time.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Closest in time.
Learning end-to-end video classification with rank-pooling
Basura Fernando and Stephen Gould · 2016
Closest in time.
Discriminative hierarchical rank pooling for activity recognition
Basura Fernando, Peter Anderson, Marcus Hutter, and Stephen Gould · 2016
Closest in time.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Closest in time.
Sparse coding and dictionary learning with linear dynamical systems
Wenbing Huang, Fuchun Sun, Lele Cao, Deli Zhao, Huaping Liu, and Mehrtash Harandi · 2016
Closest in time.
Learning cross-domain landmarks for heterogeneous domain adaptation
Yao-Hung Hubert Tsai, Yi-Ren Yeh, and Yu-Chiang Frank Wang · 2016
Closest in time.
Tensor representations via kernel linearization for action recognition from 3d skeletons
Piotr Koniusz, Anoop Cherian, and Fatih Porikli · 2016
Closest in time.
Segmental spatiotemporal cnns for fine-grained action segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D. Hager · 2016
Closest in time.
Vlad3: Encoding dynamics of deep features for action recognition
Yingwei Li, Weixin Li, Vijay Mahadevan, and Nuno Vasconcelos · 2016
Closest in time.
Spatio-temporal lstm with trust gates for 3d human action recognition
Jun Liu, Amir Shahroudy, Dong Xu, and Gang Wang · 2016
Closest in time.
Nonlinear metric learning for visual tracking
J. Lu, J. Hu, and Y. P. Tan · 2016
Closest in time.
Unsupervised Learning using Sequential Verification for Action Recognition
Ishan Misra, C. Lawrence Zitnick, and Martial Hebert · 2016
Closest in time.
Progressively parsing interactional objects for fine grained action detection
Bingbing Ni, Xiaokang Yang, and Shenghua Gao · 2016
Closest in time.
3d action recognition from novel viewpoints
Hossein Rahmani and Ajmal Mian · 2016
Closest in time.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Bharat Singh, Tim K. Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Closest in time.
Hierarchical dynamic parsing and encoding for action recognition
Bing Su, Jiahuan Zhou, Xiaoqing Ding, Hao Wang, and Ying Wu · 2016
Closest in time.
A siamese long short-term memory architecture for human re-identification
Rahul Rama Varior, Bing Shuai, Jiwen Lu, Dong Xu, and Gang Wang · 2016
Closest in time.
Long-term Temporal Convolutions for Action Recognition
Gül Varol, Ivan Laptev, and Cordelia Schmid · 2016
Closest in time.
Actions ~ transformations
Xiaolong Wang, Ali Farhadi, and Abhinav Gupta · 2016
Closest in time.