Fetching the paper…
Reading the bibliography…
We study weakly-supervised video object grounding: given a video segment and a corresponding descriptive sentence, the goal is to localize objects that are mentioned from the sentence in the video.
Weakly supervised localization and learning with generic knowledge
Thomas Deselaers, Bogdan Alexe, and Vittorio Ferrari · 2012
Earlier work this paper cites.
Learning object class detectors from weakly annotated video
Alessandro Prest, Christian Leistner, Javier Civera, Cordelia Schmid, and Vittorio Ferrari · 2012
Earlier work this paper cites.
Efficiently scaling up crowdsourced video annotation
Carl Vondrick, Donald Patterson, and Deva Ramanan · 2013
Earlier work this paper cites.
Grounded language learning from video described with sentences
Haonan Yu and Jeffrey Mark Siskind · 2013
Earlier work this paper cites.
Multi-fold mil training for weakly supervised object localization
Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid · 2014
Earlier work this paper cites.
Learning everything about anything: Webly-supervised visual concept learning
Santosh K Divvala, Ali Farhadi, and Carlos Guestrin · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei · 2014
Earlier work this paper cites.
On learning to localize objects with minimal supervision
Hyun Oh Song, Ross Girshick, Stefanie Jegelka, Julien Mairal, Zaid Harchaoui, and Trevor Darrell · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Unsupervised object discovery and tracking in video collections
Suha Kwak, Minsu Cho, Ivan Laptev, Jean Ponce, and Cordelia Schmid · 2015
Earlier work this paper cites.
Is object localization for free?-weakly-supervised learning with convolutional neural networks
Maxime Oquab, Léon Bottou, Ivan Laptev, and Josef Sivic · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Cited alongside, same era.
Segmental spatiotemporal cnns for fine-grained action segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D Hager · 2016
Cited alongside, same era.
Progressively parsing interactional objects for fine grained action detection
Bingbing Ni, Xiaokang Yang, and Shenghua Gao · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Cited alongside, same era.
Faster r-cnn: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2017
Later among the works it cites.
Generating descriptions with grounded and co-referenced people
Anna Rohrbach, Marcus Rohrbach, Siyu Tang, Seong Joon Oh, and Bernt Schiele · 2017
Later among the works it cites.
Opportunistic active learning for grounding natural language descriptions
Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Justin Hart, Peter Stone, and Raymond J Mooney · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Weakly-supervised visual grounding of phrases with linguistic structures
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Muhannad Al-Omari, Paul Duckworth, David C Hogg, and Anthony G Cohn · 2017
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2017
Cited alongside, same era.
Attend and interact: Higher-order object interactions for video understanding
Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, and Hans Peter Graf · 2017
Cited alongside, same era.
Conditional image-text embedding networks
Bryan A Plummer, Paige Kordas, M Hadi Kiapour, Shuai Zheng, Robinson Piramuthu, and Svetlana Lazebnik · 2017
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso
Cited in the paper.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong
Cited in the paper.
Sentence directed video object codiscovery
Haonan Yu and Jeffrey Mark Siskind · 2017
Later among the works it cites.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency · 2018
Closest in time.
Finding “it”: Weakly-supervised reference-aware visual grounding in instructional video
De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg, Li Fei-Fei, and Juan Carlos Niebles · 2018
Closest in time.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Closest in time.