Fetching the paper…
Reading the bibliography…
In this paper, we address a novel task, namely weakly-supervised spatio-temporally grounding natural sentence in video.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A duality based approach for realtime tv-l 1 optical flow
Christopher Zach, Thomas Pock, and Horst Bischof. 2007 · 2007
Earlier work this paper cites.
Recognition using visual phrases
Mohammad Amin Sadeghi and Ali Farhadi. 2011 · 2011
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
Cynthia Matuszek, Nicholas FitzGerald, Luke Zettlemoyer, Liefeng Bo, and Dieter Fox. 2012 · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
What are you talking about? text-to-image coreference
Chen Kong, Dahua Lin, Mohit Bansal, Raquel Urtasun, and Sanja Fidler. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Multimodal convolutional neural networks for matching image and sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. 2015 · 2015
Cited alongside, same era.
Sentence directed video object codetection
Haonan Yu and Jeffrey Mark Siskind. 2015 · 2015
Cited alongside, same era.
Netvlad: Cnn architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell. 2016 · 2016
Cited alongside, same era.
Learning to answer questions from image using convolutional neural network
A joint speakerlistener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg. 2017 · 2017
Later among the works it cites.
Discriminative bimodal networks for visual localization and detection with natural language queries
Yuting Zhang, Luyao Yuan, Yijie Guo, Zhiyuan He, I-An Huang, and Honglak Lee. 2017 · 2017
Later among the works it cites.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency. 2018 · 2018
Later among the works it cites.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia. 2018 · 2018
Later among the works it cites.
Finding “it”: Weakly-supervised reference-aware visual grounding in instructional videos
De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg, Li Fei-Fei, and Juan Carlos Niebles. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin Ma, Zhengdong Lu, and Hang Li. 2016 · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele. 2016 · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. 2017 · 2017
Cited alongside, same era.
Amc: Attention guided multi-modal correlation learning for image search
Kan Chen, Trung Bui, Chen Fang, Zhaowen Wang, and Ram Nevatia. 2017 · 2017
Cited alongside, same era.
Deep attribute-preserving metric learning for natural language object retrieval
Jianan Li, Yunchao Wei, Xiaodan Liang, Fang Zhao, Jianshu Li, Tingfa Xu, and Jiashi Feng. 2017 · 2017
Cited alongside, same era.
Generating descriptions with grounded and co-referenced people
Anna Rohrbach, Marcus Rohrbach, Siyu Tang, Seong Joon Oh, and Bernt Schiele. 2017 · 2017
Cited alongside, same era.
Multiple instance detection network with online instance classifier refinement
Peng Tang, Xinggang Wang, Xiang Bai, and Wenyu Liu. 2017 · 2017
Cited alongside, same era.
Pcl: Proposal cluster learning for weakly supervised object detection
Peng Tang, Xinggang Wang, Song Bai, Wei Shen, Xiang Bai, Wenyu Liu, and Alan Loddon Yuille. 2018 · 2018
Later among the works it cites.
Object referring in videos with language and human gaze
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool. 2018 · 2018
Later among the works it cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. 2018 · 2018
Later among the works it cites.
Grounding referring expressions in images by variational context
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang. 2018 · 2018
Later among the works it cites.
Weakly supervised phrase localization with multi-scale anchored transformer network
Fang Zhao, Jianshu Li, Jian Zhao, and Jiashi Feng. 2018 · 2018
Later among the works it cites.
Weakly-supervised video object grounding from text by loss weighting and object interaction
Luowei Zhou, Nathan Louis, and Jason J Corso. 2018 · 2018
Later among the works it cites.
Localizing natural language in videos
Jingyuan Chen, Lin Ma, Xinpeng Chen, Zequn Jie, and Jiebo Luo. 2019 · 2019
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.