Fetching the paper…
Reading the bibliography…
This paper presents a framework for localization or grounding of phrases in images using a large collection of linguistic and visual cues.
Convergence properties of the nelder-mead simplex method in low dimensions
J. C. Lagarias, J. A. Reeds, M. H. Wright, and P. E. Wright · 1998
Earlier work this paper cites.
Robust pronoun resolution with limited knowledge
R. Mitkov · 1998
Earlier work this paper cites.
Knowledge-lean coreference resolution and its relation to textual cohesion and coherence
S. Harabagiu and S. Maiorano · 1999
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
J. C. Platt · 1999
Earlier work this paper cites.
Training linear svms in linear time
T. Joachims · 2006
Earlier work this paper cites.
An empirical study of context in object detection
S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Heber · 2009
Earlier work this paper cites.
LIBSVM: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2012
Earlier work this paper cites.
A sentence is worth a thousand pixels
S. Fidler, A. Sharma, and R. Urtasun · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Parsing With Compositional Vector Grammars
R. Socher, J. Bauer, C. D. Manning, and A. Y. Ng · 2013
Earlier work this paper cites.
Selective search for object recognition
J. Uijlings, K. van de Sande, T. Gevers, and A. Smeulders · 2013
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Cited alongside, same era.
Object bank: An object-level image representation for high-level visual recognition
L.-J. Li, H. Su, Y. Lim, and L. Fei-Fei · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Scene parsing with object instances and occlusion ordering
J. Tighe, M. Niethammer, and S. Lazebnik · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Cited alongside, same era.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Visual Madlibs: Fill in the blank Image Generation and Question Answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Later among the works it cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Closest in time.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Closest in time.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Closest in time.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. Platt, L. Zitnick, and G. Zweig · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Associating neural word embeddings with deep image representations using fisher vector
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Cited alongside, same era.
Closest in time.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Closest in time.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Closest in time.
Structured matching for phrase localization
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng · 2016
Closest in time.
MSRC: Multimodal spatial regression with semantic context for phrase grounding
K. Chen, R. Kovvuri, J. Gao, and R. Nevatia · 2017
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2017
Closest in time.
ViP-CNN: Visual phrase guided convolutional neural network
Y. Li, W. Ouyang, X. Wang, and X. Tang · 2017
Closest in time.
Deep variation-structured reinforcement learning for visual relationship and attribute detection
X. Liang, L. Lee, and E. P. Xing · 2017
Closest in time.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2017
Closest in time.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Closest in time.