Fetching the paper…
Reading the bibliography…
In this paper we investigate learning visual models for the steps of ordinary tasks using weak supervision via instructional narrations and an ordered list of steps instead of strong supervision via temporal annotations.
Maximum margin clustering
L. Xu, J. Neufeld, B. Larson, and D. Schuurmans · 2004
Earlier work this paper cites.
DIFFRAC: A discriminative and flexible framework for clustering
F. Bach and Z. Harchaoui · 2007
Earlier work this paper cites.
Learning visual attributes
V. Ferrari and A. Zisserman · 2007
Earlier work this paper cites.
Visualizing data using t-sne
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Describing objects by their attributes
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth · 2009
Earlier work this paper cites.
Recognizing human actions by attributes
J. Liu, B. Kuipers, and S. Savarese · 2011
Earlier work this paper cites.
Human action recognition by learning bases of action attributes and parts
B. Yao, X. Jiang, A. Khosla, A. L. Lin, L. Guibas, and L. Fei-Fei1 · 2011
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
P. Bojanowski, R. Lajugie, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2014
Earlier work this paper cites.
You-do, i-learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video
D. Damen, T. Leelasawassuk, O. Haines, A. Calway, and W. Mayol-Cuevas · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Weakly-supervised alignment of video with text
P. Bojanowski, R. Lajugie, E. Grave, F. Bach, I. Laptev, J. Ponce, and C. Schmid · 2015
Cited alongside, same era.
What’s cookin’? Interpreting cooking videos using text, speech and vision
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy · 2015
Cited alongside, same era.
Unsupervised semantic parsing of video collections
O. Sener, A. Zamir, S. Savarese, and A. Saxena · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste Julien · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
From Red Wine to Red Tomato: Composition with Context
I. Misra, A. Gupta, and M. Hebert · 2017
Later among the works it cites.
Weakly supervised action learning with rnn based fine-to-coarse modeling
A. Richard, H. Kuehne, and J. Gall · 2017
Later among the works it cites.
Commonly uncommon: Semantic sparsity in situation recognition
M. Yatskar, V. Ordonez, L. Zettlemoyer, and A. Farhadi · 2017
Later among the works it cites.
Learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2017
Later among the works it cites.
Enriching word vectors with subword information
A. J. Piotr Bojanowski, Edouard Grave and T. Mikolov · 2017
Later among the works it cites.
Deep Clustering for Unsupervised Learning of Visual Features
M. Caron, P. Bojanowski, A. Joulin, and M. Douze · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Connectionist temporal modeling for weakly supervised action labeling
D.-A. Huang, L. Fei-Fei, and J. C. Niebles · 2016
Cited alongside, same era.
Joint discovery of object states and manipulation actions
J.-B. Alayrac, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2017
Cited alongside, same era.
Unsupervised learning by predicting noise
P. Bojanowski and A. Joulin · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Cnn architectures for large-scale audio classification
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. Weiss, , and K. Wilson · 2017
Cited alongside, same era.
Unsupervised visual-linguistic reference resolution in instructional videos
D.-A. Huang, J. J. Lim, L. Fei-Fei, and J. C. Niebles · 2017
Cited alongside, same era.
Scaling egocentric vision: The EPIC-KITCHENS dataset
D. Damen, H. Doughty, G. Maria Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Later among the works it cites.
Demo2vec: Reasoning object affordances from online videos
K. Fang, T.-L. Wu, D. Yang, S. Savarese, and J. J. Lim · 2018
Later among the works it cites.
From lifestyle vlogs to everyday interactions
D. F. Fouhey, W. Kuo, A. A. Efros, and J. Malik · 2018
Later among the works it cites.
Finding ”it”: Weakly-supervised reference-aware visual grounding in instructional video
D.-A. Huang, V. Ramanathan, D. Mahajan, L. Torresani, M. Paluri, L. Fei-Fei, and J. C. Niebles · 2018
Later among the works it cites.
Action sets: Weakly supervised action segmentation without ordering constraints
A. Richard, H. Kuehne, and J. Gall · 2018
Later among the works it cites.
Unsupervised learning and segmentation of complex activities from video
F. Sener and A. Yao · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, X. Chenliang, and J. J. Corso · 2018
Later among the works it cites.