Fetching the paper…
Reading the bibliography…
Humans refer to objects in their environments all the time, especially in dialogue with other people.
Winograd, T.: Understanding natural language. Cognitive psychology 3(1), 1–191 (1972)
1972
Earlier work this paper cites.
Grice, H.P.: Logic and conversation. In: Cole, P., Morgan, J.L. (eds.) Syntax and Semantics: Vol. 3: Speech Acts, pp. 41–58. Academic Press, San Diego, CA (1975)
1975
Earlier work this paper cites.
Jordan, P., Walker, M.: Learning attribute selections for non-pronominal expressions. In: ACL (2000)
2000
Earlier work this paper cites.
Funakoshi, K., Watanabe, S., Kuriyama, N., Tokunaga, T.: Generating referring expressions using perceptual groups. In: Natural Language Generation, pp. 51–60. Springer (2004)
2004
Earlier work this paper cites.
Brown-Schmidt, S., Tanenhaus, M.K.: Watching the eyes when talking about size: An investigation of message formulation and utterance planning. Journal of Memory and Language 54(4), 592–609 (2006)
2006
Earlier work this paper cites.
Kelleher, J.D., Kruijff, G.J.M.: Incremental generation of spatial referring expressions in situated dialog. In: ACL (2006)
2006
Earlier work this paper cites.
Viethen, J., Dale, R.: The use of spatial relations in referring expression generation. In: Proceedings of the Fifth International Natural Language Generation Conference. pp. 59–67. Association for Computational Linguistics (2008)
2008
Earlier work this paper cites.
Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., Forsyth, D.: Every picture tells a story: Generating sentences from images. In: ECCV (2010)
2010
Earlier work this paper cites.
Mitchell, M., van Deemter, K., Reiter, E.: Natural reference to objects in a visual domain. In: Proceedings of the 6th international natural language generation conference. pp. 95–104. Association for Computational Linguistics (2010)
2010
Earlier work this paper cites.
Ordonez, V., Kulkarni, G., Berg, T.L.: Im2text: Describing images using 1 million captioned photographs. In: Advances in Neural Information Processing Systems (2011)
2011
Earlier work this paper cites.
Krahmer, E., Van Deemter, K.: Computational generation of referring expressions: A survey. Computational Linguistics 38(1), 173–218 (2012)
2012
Earlier work this paper cites.
FitzGerald, N., Artzi, Y., Zettlemoyer, L.S.: Learning distributions over logical forms for referring expression generation. In: EMNLP. pp. 1914–1925 (2013)
2013
Earlier work this paper cites.
Hodosh, M., Young, P., Hockenmaier, J.: Framing image description as a ranking task: Data, models and evaluation metrics. Journal of Artificial Intelligence Research (2013)
2013
Earlier work this paper cites.
Kulkarni, G., Premraj, V., Ordonez, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., Berg, T.: Babytalk: Understanding and generating simple image descriptions. Pattern Analysis and Machine Intelligence, IEEE Transactions on (2013)
2013
Earlier work this paper cites.
Mitchell, M., Reiter, E., van Deemter, K.: Typicality and object reference. Cognitive Science (CogSci) (2013)
2013
Cited alongside, same era.
Mitchell, M., Van Deemter, K., Reiter, E.: Generating expressions that refer to visible objects. In: HLT-NAACL. pp. 1174–1184 (2013)
2013
Cited alongside, same era.
Erhan, D., Szegedy, C., Toshev, A., Anguelov, D.: Scalable object detection using deep neural networks. In: CVPR (2014)
2014
Cited alongside, same era.
Kazemzadeh, S., Ordonez, V., Matten, M., Berg, T.L.: Referitgame: Referring to objects in photographs of natural scenes. In: EMNLP. pp. 787–798 (2014)
2014
Cited alongside, same era.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV (2014)
2014
Cited alongside, same era.
Jun Zhu, X.C., Yuille, A.: Deepm: A deep part-based model for object detection and semantic part localization (2015)
2015
Later among the works it cites.
Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: CVPR (2015)
2015
Later among the works it cites.
Kiros, R., Salakhutdinov, R., Zemel, R.S.: Unifying visual-semantic embeddings with multimodal neural language models. TACL (2015)
2015
Later among the works it cites.
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S.: Ssd: Single shot multibox detector (2015)
2015
Later among the works it cites.
Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., Yuille, A.: Deep captioning with multimodal recurrent neural networks (m-rnn). ICLR (2015)
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
Socher, R., Karpathy, A., Le, Q.V., Manning, C.D., Ng, A.Y.: Grounded compositional semantics for finding and describing images with sentences. Transactions of the Association for Computational Linguistics (2014)
2014
Cited alongside, same era.
Bell, S., Zitnick, C.L., Bala, K., Girshick, R.: Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks (2015)
2015
Cited alongside, same era.
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: Long-term recurrent convolutional networks for visual recognition and description. In: CVPR (2015)
2015
Cited alongside, same era.
Fang, H., Gupta, S., Iandola, F., Srivastava, R.K., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J.C., et al.: From captions to visual concepts and back. In: CVPR (2015)
2015
Cited alongside, same era.
Girshick, R.: Fast r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1440–1448 (2015)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: NIPS (2015)
2015
Later among the works it cites.
2015
Later among the works it cites.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al.: Imagenet large scale visual recognition challenge. International Journal of Computer Vision 115(3), 211–252 (2015)
2015
Later among the works it cites.
Sadeghi, F., Zitnick, C.L., Farhadi, A.: Visalogy: Answering visual analogy questions. In: NIPS (2015)
2015
Later among the works it cites.
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator. In: CVPR (2015)
2015
Later among the works it cites.
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. ICML (2015)
2015
Later among the works it cites.
Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., Darrell, T.: Natural language object retrieval. In: CVPR (2016)
2016
Closest in time.
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A., Murphy, K.: Generation and comprehension of unambiguous object descriptions. In: CVPR (2016)
2016
Closest in time.