Fetching the paper…
Reading the bibliography…
We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described.
Logic and conversation
H. P. Grice · 1970
Earlier work this paper cites.
Understanding natural language
T. Winograd · 1972
Earlier work this paper cites.
Maximum mutual information estimation of hidden Markov model parameters for speech recognition
L. Bahl, P. Brown, P. V. de Souza, and R. Mercer · 1986
Earlier work this paper cites.
Efficient context-sensitive generation of referring expressions
E. Krahmer and M. Theune · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Building a semantically transparent corpus for the generation of referring expressions
K. van Deemter, I. van der Sluis, and A. Gatt · 2006
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgements
A. Lavie and A. Agarwal · 2007
Earlier work this paper cites.
The use of spatial relations in referring expression generation
J. Viethen and R. Dale · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The segmented and annotated IAPR TC-12 benchmark
H. J. Escalante, C. A. Hernandez, J. A. Gonzalez, A. Lopez-Lopez, M. Montes, E. F. Morales, L. E. Sucar, L. Villasenor, and M. Grubinger · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
A game-theoretic approach to generating spatial descriptions
D. Golland, P. Liang, and D. Klein · 2010
Earlier work this paper cites.
Natural reference to objects in a visual domain
M. Mitchell, K. van Deemter, and E. Reiter · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos · 2011
Earlier work this paper cites.
Computational generation of referring expressions: A survey
E. Krahmer and K. van Deemter · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Image description with a goal: Building efficient discriminating expressions for images
A. Sadovnik, Y.-I. Chiu, N. Snavely, S. Edelman, and T. Chen · 2012
Cited alongside, same era.
Learning distributions over logical forms for referring expression generation
N. FitzGerald, Y. Artzi, and L. S. Zettlemoyer · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Cited alongside, same era.
Generating expressions that refer to visible objects
M. Mitchell, K. van Deemter, and E. Reiter · 2013
Cited alongside, same era.
ImageSpirit: Verbal guided image parsing
M.-M. Cheng, S. Zheng, W.-Y. Lin, V. Vineet, P. Sturgess, N. Crook, N. J. Mitra, and P. Torr · 2014
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2015
Closest in time.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scalable object detection using deep neural networks
D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Probabilistic semantics and pragmatics: Uncertainty in language and thought
N. D. Goodman and D. Lassiter · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Cited alongside, same era.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Closest in time.
From the virtual to the real world: Referring to objects in Real-World spatial scenes
D. Gkatzia, V. Rieser, P. Bartie, and W. Mackaness · 2015
Closest in time.
Lstm: A search space odyssey
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber · 2015
Closest in time.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Closest in time.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Closest in time.
CIDEr: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, C. A. Cho, Kyunghyun, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Closest in time.