Fetching the paper…
Reading the bibliography…
We present an approach for weakly supervised learning of human actions.
Backpropagation through time: what it does and how to do it
P. J. Werbos · 1990
Earlier work this paper cites.
Using a stochastic context-free grammar as a language model for speech recognition
D. Jurafsky, C. Wooters, J. Segal, A. Stolcke, E. Fosler, G. Tajchaman, and N. Morgan · 1995
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Automatic annotation of human actions in video
O. Duchenne, I. Laptev, J. Sivic, F. Bach, and J. Ponce · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Improving the fisher kernel for large-scale image classification
F. Perronnin, J. Sánchez, and T. Mensink · 2010
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Earlier work this paper cites.
Learning latent temporal structure for complex event detection
K. Tang, L. Fei-Fei, and D. Koller · 2012
Earlier work this paper cites.
Image Classification with the Fisher Vector: Theory and Practice
J. Sanchez, F. Perronnin, T. Mensink, and J. Verbeek · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
P. Bojanowski, R. Lajugie, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. B. Arslan, and T. Serre · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Delving deeper into convolutional networks for learning video representations
N. Ballas, L. Yao, P. Chris, and A. Courville · 2016
Later among the works it cites.
Webly-supervised video recognition by mutually voting for relevant web images and web video frames
C. Gan, C. Sun, L. Duan, and B. Gong · 2016
Later among the works it cites.
Connectionist temporal modeling for weakly supervised action labeling
D.-A. Huang, L. Fei-Fei, and J. C. Niebles · 2016
Later among the works it cites.
Deep hand: How to train a cnn on 1 million hand images when your data is continuous and weakly labelled
O. Koller, H. Ney, and R. Bowden · 2016
Later among the works it cites.
An end-to-end generative framework for video segmentation and recognition
H. Kuehne, J. Gall, and T. Serre · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An empirical exploration of recurrent network architectures
R. Józefowicz, W. Zaremba, and I. Sutskever · 2015
Cited alongside, same era.
What’s cookin’? interpreting cooking videos using text, speech and vision
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy · 2015
Cited alongside, same era.
Temporal localization of fine-grained actions in videos by domain transfer from web images
C. Sun, S. Shetty, R. Sukthankar, and R. Nevatia · 2015
Cited alongside, same era.
Watch-n-patch: Unsupervised understanding of actions and relations
C. Wu, J. Zhang, S. Savarese, and A. Saxena · 2015
Cited alongside, same era.
Modeling spatial-temporal clues in a hybrid deep learning framework for video classification
Z. Wu, X. Wang, Y.-G. Jiang, H. Ye, and X. Xue · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien · 2016
Cited alongside, same era.
H. Kuehne, A. Richard, and J. Gall · 2016
Later among the works it cites.
Segmental spatiotemporal cnns for fine-grained action segmentation
C. Lea, A. Reiter, R. Vidal, and G. D. Hager · 2016
Later among the works it cites.
Shuffle and learn: Unsupervised learning using temporal order verification
I. Misra, C. L. Zitnick, and M. Hebert · 2016
Later among the works it cites.
Temporal action detection using a statistical language model
A. Richard and J. Gall · 2016
Later among the works it cites.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
B. Singh, T. K. Marks, M. Jones, O. Tuzel, and M. Shao · 2016
Later among the works it cites.
End-to-end learning of action detection from frame glimpses in videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2016
Later among the works it cites.