Fetching the paper…
Reading the bibliography…
There are substantial instructional videos on the Internet, which enables us to acquire knowledge for completing various tasks.
Wordnet: An electronic lexical database
C. Fellbaum · 1998
Earlier work this paper cites.
Optimizing the number of steps in learning tasks for complex skills
N. RJ, K. PA, and van Merriënboer JJ · 2005
Earlier work this paper cites.
A duality based approach for realtime tv-l1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li · 2009
Earlier work this paper cites.
VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method
S. E. F. de Avila, A. P. B. Lopes, A. da Luz Jr., and A. de Albuquerque Araújo · 2011
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. Zamir, and M. Shah · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
American time use survey
U. D. of Labor · 2013
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
S. Stein and S. J. McKenna · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Creating summaries from user videos
M. Gygli, H. Grabner, H. Riemenschneider, and L. J. V. Gool · 2014
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. B. Arslan, and T. Serre · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Cited alongside, same era.
Unsupervised semantic parsing of video collections
O. Sener, A. R. Zamir, S. Savarese, and A. Saxena · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Tvsum: Summarizing web videos using titles
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. D. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
J. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2016
Cited alongside, same era.
Diversity-aware multi-video summarization
R. Panda, N. C. Mithun, and A. K. Roy-Chowdhury · 2017
Later among the works it cites.
Human pose forecasting via deep markov models
S. Toyer, A. Cherian, T. Han, and S. Gould · 2017
Later among the works it cites.
R-C3D: region convolutional 3d network for temporal activity detection
H. Xu, A. Das, and K. Saenko · 2017
Later among the works it cites.
Temporal action detection with structured segment networks
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin · 2017
Later among the works it cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. Maria Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Later among the works it cites.
Weakly-supervised action segmentation with iterative soft boundary assignment
L. Ding and C. Xu · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dissimilarity-based sparse subset selection
E. Elhamifar, G. Sapiro, and S. S. Sastry · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Val Gool · 2016
Cited alongside, same era.
MSR-VTT: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Cited alongside, same era.
Video summarization with long short-term memory
K. Zhang, W. Chao, F. Sha, and K. Grauman · 2016
Cited alongside, same era.
Quo vadis, action recognition? A new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
PKU-MMD: A large scale benchmark for continuous multi-modal human action understanding
L. Chunhui, H. Yueyu, L. Yanghao, S. Sijie, and L. Jiaying · 2017
Cited alongside, same era.
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2018
Later among the works it cites.
Finding "it": Weakly-supervised reference-aware visual grounding in instructional videos
D.-A. Huang, S. Buch, L. Dery, A. Garg, L. Fei-Fei, and J. Carlos Niebles · 2018
Later among the works it cites.
Action sets: Weakly supervised action segmentation without ordering constraints
A. Richard, H. Kuehne, and J. Gall · 2018
Later among the works it cites.
Neuralnetwork-viterbi: A framework for weakly supervised video learning
A. Richard, H. Kuehne, A. Iqbal, and J. Gall · 2018
Later among the works it cites.
Fine-grained video captioning for sports narrative
H. Yu, S. Cheng, B. Ni, M. Wang, J. Zhang, and X. Yang · 2018
Later among the works it cites.
Weakly-supervised video object grounding from text by loss weighting and object interaction
L. Zhou, N. Louis, and J. J. Corso · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Later among the works it cites.
End-to-end dense video captioning with masked transformer
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong · 2018
Later among the works it cites.