Fetching the paper…
Reading the bibliography…
Thanks to the success of object detection technology, we can retrieve objects of the specified classes even from huge image collections.
Wordnet: a lexical database for English
George A. Miller. 1995 · 1995
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
Pedro F Felzenszwalb, Ross B Girshick, David McAllester, Deva Ramanan, and David Forsyth. 2010 · 2010
Earlier work this paper cites.
Product quantization for nearest neighbor search
Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011 · 2011
Earlier work this paper cites.
Multiple queries for large scale specific object retrieval
Relja Arandjelovi, Andrew Zisserman, Relja Arandjelovic, Andrew Zisserman, Relja Arandjelovi, Andrew Zisserman, Relja Arandjelovic, and Andrew Zisserman. 2012 · 2012
Earlier work this paper cites.
Object retrieval and localization with spatially-constrained similarity measure and k-nn re-ranking
Xiaohui Shen, Zhe Lin, Jonathan Brandt, Shai Avidan, and Ying Wu. 2012 · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Immediate, scalable object category detection
Yusuf Aytar and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, U C Berkeley, and Jitendra Malik. 2014 · 2014
Earlier work this paper cites.
Open-vocabulary object retrieval
Sergio Guadarrama, Erik Rodner, Kate Saenko, Ning Zhang, Ryan Farrell, Jeff Donahue, and Trevor Darrell. 2014 · 2014
Earlier work this paper cites.
Referitgame: referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L Berg. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. 2014 · 2014
Earlier work this paper cites.
Locality in generic instance search from one example
Ran Tao, Efstratios Gavves, Cees G M Snoek, and Arnold W M Smeulders. 2014 · 2014
Earlier work this paper cites.
On-the-fly learning for visual search of large-scale image and video datasets
Ken Chatfield, Relja Arandjelovi, Andrew Zisserman, Relja Arandjelović, Omkar Parkhi, and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Adam: a method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf. 2015 · 2015
Cited alongside, same era.
Flickr30k entities: collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Cited alongside, same era.
Faster r-cnn: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Densecap: fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Later among the works it cites.
Visual genome: connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalanditis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, Li Fei-Fei, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. 2016 · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy. 2016 · 2016
Later among the works it cites.
Image question answering using convolutional neural network with dynamic parameter prediction
Hyeonwoo Noh, Paul Hongsuck Seo, and Bohyung Han. 2016 · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
Jimmy Lei Ba, Kevin Swersky, Sanja Fidler, and Ruslan Salakhutdinov. 2016 · 2016
Cited alongside, same era.
Dynamic filter networks
Bert De Brabandere, Xu Jia, Tinne Tuytelaars, and Luc Van Gool. 2016 · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Cited alongside, same era.
Large-scale r-cnn with classifier adaptive quantization
Ryota Hinami and Shin’ichi Satoh. 2016 · 2016
Cited alongside, same era.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell. 2016 · 2016
Cited alongside, same era.
Phrase localization and visual relationship Detection with Comprehensive Image-Language Cues
Bryan A Plummer, Christopher M Cervantes, and C V Aug. 2017a
Cited in the paper.
Training region-based object detectors with online hard example mining
Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. 2016 · 2016
Later among the works it cites.
Particular object retrieval with integral max-pooling of cnn activations
Giorgos Tolias, Ronan Sicre, and Hervé Jégou. 2016 · 2016
Later among the works it cites.
Query-guided Regression Network with Context Policy for Phrase Grounding
Kan Chen, Rama Kovvuri, and Ram Nevatia. 2017 · 2017
Closest in time.
Region-based image retrieval revisited
Ryota Hinami, Yusuke Matsui, and Shin’ichi Satoh. 2017 · 2017
Closest in time.
YOLO9000: better, faster, stronger
Joseph Redmon and Ali Farhadi. 2017 · 2017
Closest in time.
Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries
Yuting Zhang, Luyao Yuan, Yijie Guo, Zhiyuan He, I-An Huang, and Honglak Lee. 2017 · 2017
Closest in time.