Fetching the paper…
Reading the bibliography…
Referring expressions are natural language constructions used to identify particular objects within a scene.
Understanding natural language
T. Winograd · 1972
Earlier work this paper cites.
Logic and conversation
H. P. Grice · 1975
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
M. Grubinger, P. Clough, H. Müller, and T. Deselaers · 2006
Earlier work this paper cites.
A game-theoretic approach to generating spatial descriptions
D. Golland, P. Liang, and D. Klein · 2010
Earlier work this paper cites.
Computational generation of referring expressions: A survey
E. Krahmer and K. Van Deemter · 2012
Earlier work this paper cites.
Learning distributions over logical forms for referring expression generation
N. FitzGerald, Y. Artzi, and L. S. Zettlemoyer · 2013
Earlier work this paper cites.
Generating expressions that refer to visible objects
M. Mitchell, K. Van Deemter, and E. Reiter · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Embodied collaborative referring expression generation in situated human-robot interaction
R. Fang, M. Doering, and J. Y. Chai · 2015
Cited alongside, same era.
Guiding the long-short term memory model for image caption generation
X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars · 2015
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Closest in time.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Closest in time.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Closest in time.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Closest in time.
Modeling context between objects for referring expression understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank description generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Cited alongside, same era.
Reasoning about pragmatics with neural listeners and speakers
J. Andreas and D. Klein · 2016
Cited alongside, same era.
Interpreting multimodal referring expressions in real time
M. Eldon, D. Whitney, and S. Tellex · 2016
Cited alongside, same era.
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Closest in time.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Closest in time.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Closest in time.
Structured matching for phrase localization
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng · 2016
Closest in time.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Closest in time.