Fetching the paper…
Reading the bibliography…
We propose an unsupervised method for reference resolution in instructional videos, where the goal is to temporally link an entity (e.g., "dressing") to the action (e.g., "mix yogurt") that produced it.
Names and faces in the news
T. L. Berg, A. C. Berg, J. Edwards, M. Maire, R. White, Y.-W. Teh, E. Learned-Miller, and D. A. Forsyth · 2004
Earlier work this paper cites.
Automatic annotation of human actions in video
O. Duchenne, I. Laptev, J. Sivic, F. Bach, and J. Ponce · 2009
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. Shi, and L. S. Davis · 2009
Earlier work this paper cites.
Interpreting written how-to instructions
T. A. Lau, C. Drews, and J. Nichols · 2009
Earlier work this paper cites.
An efficient sparse metric learning in high-dimensional space via l 1-penalized log-determinant regularization
G.-J. Qi, J. Tang, Z.-J. Zha, T.-S. Chua, and H.-J. Zhang · 2009
Earlier work this paper cites.
Incremental reference resolution: The task, metrics for evaluation, and a bayesian filtering model that is sensitive to disfluencies
D. Schlangen, T. Baumann, and M. Atterer · 2009
Earlier work this paper cites.
Cross-caption coreference resolution for automatic image understanding
M. Hodosh, P. Young, C. Rashtchian, and J. Hockenmaier · 2010
Earlier work this paper cites.
Toward understanding natural language directions
T. Kollar, S. Tellex, D. Roy, and N. Roy · 2010
Earlier work this paper cites.
Stanford’s multi-pass sieve coreference resolution system at the conll-2011 shared task
H. Lee, Y. Peirsman, A. Chang, N. Chambers, M. Surdeanu, and D. Jurafsky · 2011
Earlier work this paper cites.
Finding actors and actions in movies
P. Bojanowski, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2013
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
Easy victories and uphill battles in coreference resolution
G. Durrett and D. Klein · 2013
Earlier work this paper cites.
A sentence is worth a thousand pixels
S. Fidler, A. Sharma, and R. Urtasun · 2013
Earlier work this paper cites.
Jointly learning to parse and perceive: Connecting natural language to the physical world
J. Krishnamurthy and T. Kollar · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Video event understanding using natural language descriptions
V. Ramanathan, P. Liang, and L. Fei-Fei · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Earlier work this paper cites.
Learning structured perceptrons for coreference resolution with latent antecedents and non-local features
A. Björkelund and J. Kuhn · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Earlier work this paper cites.
Visual semantic search: Retrieving videos via complex textual queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Earlier work this paper cites.
Cooking with semantics
J. Malmaud, E. J. Wagner, N. Chang, and K. Murphy · 2014
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
C. D. Manning, M. Surdeanu, J. Bauer, J. R. Finkel, S. Bethard, and D. McClosky · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
Parsing videos of actions with segmental grammars
H. Pirsiavash and D. Ramanan · 2014
Cited alongside, same era.
Linking people in videos with “their” names using coreference resolution
V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Cited alongside, same era.
Learning perceptually grounded word meanings from unaligned parallel data
S. Tellex, P. Thaker, J. Joseph, and N. Roy · 2014
Unsupervised semantic parsing of video collections
O. Sener, A. R. Zamir, S. Savarese, and A. Saxena · 2015
Later among the works it cites.
Generating notifications for missing actions: Don’t forget to turn the lights off!
B. Soran, A. Farhadi, and L. Shapiro · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Later among the works it cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Later among the works it cites.
Sort story: Sorting jumbled images and captions into stories
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Instructional videos for unsupervised harvesting and learning of action examples
S.-I. Yu, L. Jiang, and A. Hauptmann · 2014
Cited alongside, same era.
Alignment-based compositional semantics for instruction following
J. Andreas and D. Klein · 2015
Cited alongside, same era.
Weakly-supervised alignment of video with text
P. Bojanowski, R. Lajugie, E. Grave, F. Bach, I. Laptev, J. Ponce, and C. Schmid · 2015
Cited alongside, same era.
Predicting the structure of cooking recipes
J. Jermsurawong and N. Habash · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
H. Agrawal, A. Chandrasekaran, D. Batra, D. Parikh, and M. Bansal · 2016
Later among the works it cites.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien · 2016
Later among the works it cites.
Structure inference machines: Recurrent neural networks for analyzing relations in group activity recognition
Z. Deng, A. Vahdat, H. Hu, and G. Mori · 2016
Later among the works it cites.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2016
Later among the works it cites.
Visual storytelling
T.-H. K. Huang, F. Ferraro, N. Mostafazadeh, I. Misra, A. Agrawal, J. Devlin, R. Girshick, X. He, P. Kohli, D. Batra, et al · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Jointly learning grounded task structures from language instruction and visual demonstration
C. Liu, S. Yang, S. Saba-Sadiya, N. Shukla, Y. He, S.-C. Zhu, and J. Y. Chai · 2016
Later among the works it cites.
Simpler context-dependent logical forms via model projections
R. Long, P. Pasupat, and P. Liang · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Later among the works it cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Learning visual storylines with skipping recurrent neural networks
G. A. Sigurdsson, X. Chen, and A. Gupta · 2016
Later among the works it cites.
Robot learning with a spatial, temporal, and causal and-or graph
C. Xiong, N. Shukla, W. Xiong, and S.-C. Zhu · 2016
Later among the works it cites.
Grounded semantic role labeling
S. Yang, Q. Gao, C. Liu, C. Xiong, S.-C. Zhu, and J. Y. Chai · 2016
Later among the works it cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Later among the works it cites.
Joint discovery of object states and manipulating actions
J.-B. Alayrac, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2017
Closest in time.