Fetching the paper…
Reading the bibliography…
Localizing moments in a longer video via natural language queries is a new, challenging task at the intersection of language and video understanding.
Phrase localization and visual relationship detection with comprehensive image-language cues
Bryan A Plummer, Arun Mallya, Christopher M Cervantes, Julia Hockenmaier, and Svetlana Lazebnik. 2017 · 1937
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Grounding the lexical semantics of verbs in visual perception using force dynamics and event logic
Jeffrey Mark Siskind. 2001 · 2001
Earlier work this paper cites.
Temporal prepositions and their logic
Ian Pratt-Hartmann. 2004 · 2004
Earlier work this paper cites.
The preposition project
Kenneth C Litkowski and Orin Hargraves. 2005 · 2005
Earlier work this paper cites.
An interval logic for natural language semantics
Savas Konur. 2008 · 2008
Earlier work this paper cites.
Parsing time: Learning to interpret time expressions
Gabor Angeli, Christopher D Manning, and Daniel Jurafsky. 2012 · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. 2014 · 2014
Earlier work this paper cites.
Cooking with semantics
Jon Malmaud, Earl J. Wagner, Nancy Chang, and Kevin Murphy. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Cited alongside, same era.
Mise en place: Unsupervised interpretation of instructional recipes
C. Kiddon, G. T. Ponnuraj, L. Zettlemoyer, and Y. Choi. 2015 · 2015
Cited alongside, same era.
What’s cookin’? Interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nick Johnston, Andrew Rabinovich, and Kevin Murphy. 2015 · 2015
Cited alongside, same era.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. 2016 · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016 · 2016
Later among the works it cites.
Detecting visual relationships with deep relational networks
Bo Dai, Yuqi Zhang, and Dahua Lin. 2017 · 2017
Later among the works it cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Later among the works it cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Cited alongside, same era.
A compositional framework for grounding language inference, generation, and acquisition in video
Haonan Yu, N Siddharth, Andrei Barbu, and Jeffrey Mark Siskind. 2015 · 2015
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2016 · 2016
Cited alongside, same era.
Marioqa: Answering questions by watching gameplay videos
Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2016 · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2017 · 2017
Later among the works it cites.
Unsupervised visual-linguistic reference resolution in instructional videos
De-An Huang, Joseph J. Lim, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Later among the works it cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Later among the works it cites.
Weakly-supervised learning of visual relations
Julia Peyre, Ivan Laptev, Cordelia Schmid, and Josef Sivic. 2017 · 2017
Later among the works it cites.
Time expression analysis and recognition using syntactic token types and general heuristic rules
Xiaoshi Zhong, Aixin Sun, and Erik Cambria. 2017 · 2017
Later among the works it cites.