Fetching the paper…
Reading the bibliography…
Classifying videos according to content semantics is an important problem with a wide range of applications.
A combined corner and edge detector
C. Harris and M. J. Stephens · 1988
Earlier work this paper cites.
Multiple kernel learning, conic duality, and the smo algorithm
F. R. Bach, G. R. Lanckriet, and M. I. Jordan · 2004
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
D. G. Lowe · 2004
Earlier work this paper cites.
Framewise phoneme classification with bidirectional lstm and other neural network architectures
A. Graves and J. Schmidhuber · 2005
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2007
Earlier work this paper cites.
Conditional random fields for activity recognition
D. L. Vail, M. M. Veloso, and J. D. Lafferty · 2007
Earlier work this paper cites.
J. Zhang, M. Marszałek, S. Lazebnik, and C. Schmid · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Klaser, M. Marszałek, and C. Schmid · 2008
Earlier work this paper cites.
Expandable data-driven graphical modeling of human actions based on salient postures
W. Li, Z. Zhang, and Z. Liu · 2008
Earlier work this paper cites.
Short-term audio-visual atoms for generic video concept classification
W. Jiang, C. Cotton, S.-F. Chang, D. Ellis, and A. Loui · 2009
Earlier work this paper cites.
Multiple kernels for object detection
A. Vedaldi, V. Gulshan, M. Varma, and A. Zisserman · 2009
Earlier work this paper cites.
Evaluation of local spatio-temporal features for action recognition
H. Wang, M. M. Ullah, A. Klaser, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Max-margin hidden conditional random fields for human action recognition
Y. Wang and G. Mori · 2009
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2010
Earlier work this paper cites.
Knowledge based activity recognition with dynamic bayesian network
Z. Zeng and Q. Ji · 2010
Earlier work this paper cites.
Consumer video understanding: A benchmark database and an evaluation of human and machine performance
Y.-G. Jiang, G. Ye, S.-F. Chang, D. Ellis, and A. C. Loui · 2011
Earlier work this paper cites.
Lp-norm multiple kernel learning
M. Kloft, U. Brefeld, S. Sonnenburg, and A. Zien · 2011
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Ng · 2011
Earlier work this paper cites.
Three things everyone should know to improve object retrieval
R. Arandjelovic and A. Zisserman · 2012
Cited alongside, same era.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Cited alongside, same era.
Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition
G. E. Dahl, D. Yu, L. Deng, and A. Acero · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
N. Srivastava and R. Salakhutdinov · 2012
Cited alongside, same era.
Discovering joint audio-visual codewords for video event detection
I.-H. Jhuo, G. Ye, S. Gao, D. Liu, Y.-G. Jiang, D. T. Lee, and S.-F. Chang · 2014
Later among the works it cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Later among the works it cites.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Later among the works it cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Later among the works it cites.
Beyond gaussian pyramid: Multi-skip feature stacking for action recognition
Z. Lan, M. Lin, X. Li, A. G. Hauptmann, and B. Raj · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning latent temporal structure for complex event detection
K. Tang, L. Fei-Fei, and D. Koller · 2012
Cited alongside, same era.
Robust late fusion with rank minimization
G. Ye, D. Liu, I.-H. Jhuo, and S.-F. Chang · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. E. Hinton · 2013
Cited alongside, same era.
Sample-specific late fusion for visual category recognition
D. Liu, K.-T. Lai, G. Ye, M.-S. Chen, and S.-F. Chang · 2013
Cited alongside, same era.
Action and event recognition with fisher vectors on a compact feature set
D. Oneata, J. Verbeek, C. Schmid, et al · 2013
Cited alongside, same era.
Image classification with the fisher vector: Theory and practice
J. Sánchez, F. Perronnin, T. Mensink, and J. Verbeek · 2013
Cited alongside, same era.
A. J. Ma and P. C. Yuen · 2014
Later among the works it cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Later among the works it cites.
CNN features off-the-shelf: an astounding baseline for recognition
A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Later among the works it cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Going Deeper with Convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
C3d: Generic features for video analysis
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2014
Later among the works it cites.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2014
Later among the works it cites.
Exploring inter-feature and inter-class relationships with deep neural networks for video classification
Z. Wu, Y.-G. Jiang, J. Wang, J. Pu, and X. Xue · 2014
Later among the works it cites.
Gated feedback recurrent neural networks
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio · 2015
Closest in time.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Exploiting image-trained cnn architectures for unconstrained video classification
S. Zha, F. Luisier, W. Andrews, N. Srivastava, and R. Salakhutdinov · 2015
Closest in time.