Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
THUMOS challenge: Action recognition with a large number of classes, 2014
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Cited alongside, same era.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. Arslan, and T. Serre · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Cited alongside, same era.
What’s cookin’? Interpreting cooking videos using text, speech and vision
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy · 2015
Cited alongside, same era.
Unsupervised semantic parsing of video collections
O. Sener, A. Zamir, S. Savarese, and A. Saxena · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien · 2016
Cited alongside, same era.
Connectionist temporal modeling for weakly supervised action labeling
D.-A. Huang, L. Fei-Fei, and J. C. Niebles · 2016
Cited alongside, same era.
Temporal tessellation for video annotation and summarization
Original
D. Kaufman, G. Levi, T. Hassner, and L. Wolf · 2016
Cited alongside, same era.
Weakly supervised learning of actions from transcripts
Original
H. Kuehne, A. Richard, and J. Gall · 2016
Cited alongside, same era.