Fetching the paper…
Reading the bibliography…
We address the problem of grounding free-form textual phrases by using weak supervision from image-caption pairs.
Adaptive distance metric learning for clustering
Jieping Ye, Zheng Zhao, and Huan Liu · 2007
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al · 2013
Earlier work this paper cites.
Weakly supervised object detection with posterior regularization
Hakan Bilen, Marco Pedersoli, and Tinne Tuytelaars · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Gated feedback recurrent neural networks
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Multiple instance learning for soft bags via top instances
Weixin Li and Nuno Vasconcelos · 2015
Earlier work this paper cites.
Multimodal convolutional neural networks for matching image and sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2015
Cited alongside, same era.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Later among the works it cites.
Hierarchical multimodal lstm for dense visual-semantic embedding
Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua · 2017
Later among the works it cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Later among the works it cites.
Weakly-supervised visual grounding of phrases with linguistic structures
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee · 2017
Later among the works it cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan R Salakhutdinov, and Alexander J Smola · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Weakly supervised object localization with multi-fold multiple instance learning
Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid · 2017
Cited alongside, same era.
Linking image and text with 2-way nets
Aviv Eisenschtat and Lior Wolf · 2017
Cited alongside, same era.
Vse++: improved visual-semantic embeddings
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Cited alongside, same era.
Karuna Ahuja, Karan Sikka, Anirban Roy, and Ajay Divakaran · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Later among the works it cites.
Zero-shot object detection
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran · 2018
Later among the works it cites.
Knowledge aided consistency for weakly supervised phrase grounding
Kan Chen, Jiyang Gao, and Ram Nevatia · 2018
Later among the works it cites.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency · 2018
Later among the works it cites.
Finding beans in burgers: Deep semantic-visual embedding with localization
Martin Engilberge, Louis Chevallier, Patrick Pérez, and Matthieu Cord · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik · 2018
Later among the works it cites.
Top-down neural attention by excitation backprop
Jianming Zhang, Sarah Adel Bargal, Zhe Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff · 2018
Later among the works it cites.
Multi-level multimodal common semantic space for image-phrase grounding
Hassan Akbari, Svebor Karaman, Surabhi Bhargava, Brian Chen, Carl Vondrick, and Shih-Fu Chang · 2019
Closest in time.