Fetching the paper…
Reading the bibliography…
Given a natural language query, a phrase grounding system aims to localize mentioned objects in an image.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Weakly supervised learning of part-based spatial models for visual object recognition
D. J. Crandall and D. P. Huttenlocher · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li · 2009
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Weakly supervised learning of interactions between humans and objects
A. Prest, C. Schmid, and V. Ferrari · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. Van D. S., T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
Open-vocabulary object retrieval
S. Guadarrama, E. Rodner, K. Saenko, N. Zhang, R. Farrell, J. Donahue, and T. Darrell · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F.-F. Li · 2014
Earlier work this paper cites.
Referit game: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
C. D. Manning, M. Surdeanu, J. Bauer, J. R. Finkel, S. Bethard, and D. McClosky · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
K. Andrej and F.-F. Li · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Fast R-CNN
R. Girshick · 2015
Leveraging visual question answering for image-caption ranking
X. Lin and D. Parikh · 2016
Later among the works it cites.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Later among the works it cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
AMC: Attention guided multi-modal correlation learning for image search
K. Chen, T. Bui, C. Fang, Z. Wang, and R. Nevatia · 2017
Later among the works it cites.
MSRC: Multimodal spatial regression with semantic context for phrase grounding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Is object localization for free?-weakly-supervised learning with convolutional neural networks
M. Oquab, L. Bottou, I. Laptev, and J. Sivic · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
ABC-CNN: An attention based convolutional neural network for visual question answering
K. Chen, J. Wang, L.-C. Chen, H. Gao, W. Xu, and R. Nevatia · 2016
Cited alongside, same era.
K. Chen, R. Kovvuri, J. Gao, and R. Nevatia · 2017
Later among the works it cites.
Query-guided regression network with context policy for phrase grounding
K. Chen, R. Kovvuri, and R. Nevatia · 2017
Later among the works it cites.
Stylenet: Generating attractive visual captions with styles
C. Gan, Z. Gan, X. He, J. Gao, and L. Deng · 2017
Later among the works it cites.
VQS: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
C. Gan, Y. Li, H. Li, C. Sun, and B. Gong · 2017
Later among the works it cites.
TALL: Temporal activity localization via language query
J. Gao, C. Sun, Z. Yang, and R. Nevatia · 2017
Later among the works it cites.
Weakly-supervised visual grounding of phrases with linguistic structures
F. Xiao, L. Sigal, and Y. J. Lee · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Later among the works it cites.
Motion-appearance co-memory networks for video question answering
J. Gao, R. Ge, K. Chen, and R. Nevatia · 2018
Closest in time.