Fetching the paper…
Reading the bibliography…
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing.
R. Gerber and H.-H. Nagel, “Knowledge representation for the generation of quantified natural language descriptions of vehicle traffic in image sequences,” in ICIP . IEEE, 1996
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, 1997
1997
Earlier work this paper cites.
X. Chen and C. L. Zitnick, “Mind?s eye: A recurrent visual representation for image caption generation,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W. J. Zhu, “BLEU: A method for automatic evaluation of machine translation,” in ACL , 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out: Proceedings of the ACL-04 workshop , vol. 8, 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , vol. 29, 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in ECCV , 2010
2010
Earlier work this paper cites.
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu, “I2t: Image parsing to text description,” Proceedings of the IEEE , vol. 98, no. 8, 2010
2010
Earlier work this paper cites.
A. Aker and R. Gaizauskas, “Generating image descriptions using dependency relational patterns,” in ACL , 2010
2010
Earlier work this paper cites.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using amazon’s mechanical turk,” in NAACL HLT Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk , 2010, pp. 139–147
2010
Earlier work this paper cites.
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “Baby talk: Understanding and generating simple image descriptions,” in CVPR , 2011
2011
Earlier work this paper cites.
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi, “Composing simple image descriptions using web-scale n-grams,” in Conference on Computational Natural Language Learning , 2011
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. L. Berg, “Im2text: Describing images using 1 million captioned photographs,” in NIPS , 2011
2011
Earlier work this paper cites.
L. Breiman, “Bagging predictors,” Machine Learning , vol. 24, pp. 123–140, 1996
2011
Earlier work this paper cites.
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. C. Berg, K. Yamaguchi, T. L. Berg, K. Stratos, and H. D. III, “Midge: Generating image descriptions from computer vision detections,” in EACL , 2012
2012
Earlier work this paper cites.
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi, “Collective generation of natural image descriptions,” in ACL , 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
D. Elliott and F. Keller, “Image description using visual dependency representations.” in EMNLP , 2013
2013
Cited alongside, same era.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics.” JAIR , vol. 47, 2013
2013
Cited alongside, same era.
R. Kiros and R. Z. R. Salakhutdinov, “Multimodal neural language models,” in NIPS Deep Learning Workshop , 2013
2013
Cited alongside, same era.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in ICLR , 2013
2013
Cited alongside, same era.
A. Graves, “Generating sequences with recurrent neural networks,” arXiv:1308.0850 , 2013
2013
Cited alongside, same era.
2014
Later among the works it cites.
2014
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” 2014
2014
Cited alongside, same era.
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in EMNLP , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NIPS , 2014
2014
Cited alongside, same era.
P. Kuznetsova, V. Ordonez, T. Berg, and Y. Choi, “Treetalk: Composition and compression of trees for image descriptions,” ACL , vol. 2, no. 10, 2014
2014
Cited alongside, same era.
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik, “Improving image-sentence embeddings using large weakly annotated photo collections,” in ECCV , 2014
2014
Cited alongside, same era.
R. Socher, A. Karpathy, Q. V. Le, C. Manning, and A. Y. Ng, “Grounded compositional semantics for finding and describing images with sentences,” in ACL , 2014
2014
Cited alongside, same era.
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. Platt, C. L. Zitnick, and G. Zweig, “From captions to visual concepts and back,” in CVPR , 2015
2015
Later among the works it cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Later among the works it cites.
——, “Deep captioning with multimodal recurrent neural networks (m-rnn),” ICLR , 2015
2015
Later among the works it cites.
R. Kiros, R. Salakhutdinov, and R. S. Zemel, “Unifying visual-semantic embeddings with multimodal neural language models,” in Transactions of the Association for Computational Linguistics , 2015
2015
Later among the works it cites.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrel, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015
2015
Later among the works it cites.
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” ICML , 2015
2015
Later among the works it cites.
A. Karpathy and F.-F. Li, “Deep visual-semantic alignments for generating image descriptions,” in CVPR , 2015
2015
Later among the works it cites.
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell, “Language models for image captioning: The quirks and what works,” in ACL , 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in CVPR , 2015
2015
Later among the works it cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,” Nov. 2015. [Online]. Available: http://download.tensorflow.org/paper/whitepaper2015.pdf
2015
Later among the works it cites.
S. Bengio, O. Vinyals, N. Jaitly, and N.Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Advances in Neural Information Processing Systems, NIPS , 2015
2015
Later among the works it cites.