Fetching the paper…
Reading the bibliography…
We present a novel method for aligning a sequence of instructions to a video of someone carrying out a task.
The story picturing Engine-A system for automatic text illustration
Joshi, D., Wang, J. Z., and Li, J. (2006) · 2006
Earlier work this paper cites.
Learning accurate, compact, and interpretable tree annotation
Petrov, S., Barrett, L., Thibaux, R., and Klein, D. (2006) · 2006
Earlier work this paper cites.
YAGO: A Large Ontology from Wikipedia and WordNet
Suchanek, F. M., Kasneci, G., and Weikum, G. (2007) · 2007
Earlier work this paper cites.
Toward an architecture for never-ending language learning
Carlson, A., Betteridge, J., Kisiel, B., Settles, B., Jr., E. H., and Mitchell, T. (2010) · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. L. and Dolan, W. B. (2011) · 2011
Earlier work this paper cites.
Open Information Extraction: the Second Generation
Etzioni, O., Fader, A., Christensen, J., Soderland, S., and Mausam (2011) · 2011
Earlier work this paper cites.
Spice it up?: Mining refinements to online instructions from user generated content
Druck, G. and Pang, B. (2012) · 2012
Earlier work this paper cites.
YouTube2Text: Recognizing and describing arbitrary activities using semantic hierarchies and Zero-Shot recognition
Guadarrama, S., Krishnamoorthy, N., Malkarnenkar, G., Venugopalan, S., Mooney, R., Darrell, T., and Saenko, K. (2013) · 2013
Cited alongside, same era.
Large scale deep neural network acoustic modeling with semi-supervised training data for YouTube video transcription
Liao, H., McDermott, E., and Senior, A. (2013) · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Cited alongside, same era.
Grounded language learning from video described with sentences
Yu, H. and Siskind, J. (2013) · 2013
Cited alongside, same era.
Food-101 – mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L. (2014) · 2014
Cited alongside, same era.
Unsupervised alignment of natural language instructions with video segments
Naim, I., Song, Y. C., Liu, Q., Kautz, H., Luo, J., and Gildea, D. (2014) · 2014
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2014) · 2014
Later among the works it cites.
RoboBrain: Large-Scale knowledge engine for robots
Saxena, A., Jain, A., Sener, O., Jami, A., Misra, D. K., and Koppula, H. S. (2014) · 2014
Later among the works it cites.
Very deep convolutional networks for Large-Scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Later among the works it cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge vault: A web-scale approach to probabilistic knowledge fusion
Dong, X., Gabrilovich, E., Heitz, G., Horn, W., Lao, N., Murphy, K., Strohmann, T., Sun, S., and Zhang, W. (2014) · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T. (2014) · 2014
Cited alongside, same era.
A database for fine grained activity detection of cooking activities
Rohrbach, M., Amin, S., Andriluka, M., and Schiele, B. (2012a)
Cited in the paper.
Script data for Attribute-Based recognition of composite activities
Rohrbach, M., Regneri, M., Andriluka, M., Amin, S., Pinkal, M., and Schiele, B. (2012b)
Cited in the paper.
Thomason, J., Venugopalan, S., Guadarrama, S., Saenko, K., and Mooney, R. (2014) · 2014
Later among the works it cites.
Instructional videos for unsupervised harvesting and learning of action examples
Yu, S.-I., Jiang, L., and Hauptmann, A. (2014) · 2014
Later among the works it cites.
Robot learning manipulation action plans by “watching” unconstrained videos from the world wide web
Yang, Y., Li, Y., Fermüller, C., and Aloimonos, Y. (2015) · 2015
Closest in time.