Fetching the paper…
Reading the bibliography…
Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Association for Computational Linguistics (ACL) . Association for Computational Linguistics, 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in The 42nd Annual Meeting of the Association for Computational Linguistics (ACL) Workshop , 2004, p. 10
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in The 43rd Annual Meeting of the Association for Computational Linguistics (ACL) Workshop , 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in European Conference on Computer Vision (ECCV) . Springer, 2010, pp. 15–29
2010
Earlier work this paper cites.
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos, “Corpus-guided sentence generation of natural images,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2011, pp. 444–454
2011
Earlier work this paper cites.
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III, “Midge: Generating image descriptions from computer vision detections,” in Association for Computational Linguistics (ACL) . Association for Computational Linguistics, 2012, pp. 747–756
2012
Earlier work this paper cites.
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “Babytalk: Understanding and generating simple image descriptions,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 35, no. 12, pp. 2891–2903, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Z. Yu, F. Wu, Y. Yang, Q. Tian, J. Luo, and Y. Zhuang, “Discriminative coupled dictionary hashing for fast cross-media retrieval,” in Association for Computing Machinery’s Special Interest Group on Information Retrieval (ACM SIGIR)) , 2014, pp. 395–404
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in Neural Information Processing Systems (NIPS) , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
A. Karpathy, A. Joulin, and L. F. Fei-Fei, “Deep fragment embeddings for bidirectional image sentence mapping,” in Advances in Neural Information Processing Systems (NIPS) , 2014, pp. 1889–1897
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research (JMLR) , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. Manning, “Glove: Global vectors for word representation,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
J. Yu, Y. Rui, Y. Y. Tang, and D. Tao, “High-order distance-based multiview stochastic learning in image classification,” IEEE Transactions On Cybernetics (CYB) , vol. 44, no. 12, pp. 2431–2442, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 4566–4575
2015
Cited alongside, same era.
2015
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 1–9
2015
Cited alongside, same era.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 2625–2634
2015
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 7008–7024
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NIPS) , 2017, pp. 6000–6010
2017
Later among the works it cites.
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei, “Boosting image captioning with attributes,” in IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 4894–4902
2017
Later among the works it cites.
L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T.-S. Chua, “Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 5659–5667
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 3128–3137
2015
Cited alongside, same era.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems (NIPS) , 2015, pp. 91–99
2015
Cited alongside, same era.
R. Girshick, “Fast r-cnn,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 1440–1448
2015
Cited alongside, same era.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 3156–3164
2015
Cited alongside, same era.
2016
Cited alongside, same era.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” in European Conference on Computer Vision (ECCV) , 2016, pp. 451–466
2016
Cited alongside, same era.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in Advances in neural information processing systems (NIPS) , 2016, pp. 289–297
2016
Cited alongside, same era.
2016
Cited alongside, same era.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision (IJCV) , vol. 123, no. 1, pp. 32–73, 2017
2017
Later among the works it cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in IEEE international conference on computer vision (ICCV) , 2017, pp. 2980–2988
2017
Later among the works it cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 1492–1500
2017
Later among the works it cites.
J.-H. Kim, J. Jun, and B.-T. Zhang, “Bilinear attention networks,” Advances in Neural Information Processing Systems (NIPS) , 2018
2018
Later among the works it cites.
Z. Yu, J. Yu, C. Xiang, Z. Zhao, Q. Tian, and D. Tao, “Rethinking diversified and discriminative proposal generation for visual grounding,” International Joint Conference on Artificial Intelligence (IJCAI) , pp. 1114–1120, 2018
2018
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Later among the works it cites.
N. Xu, A.-A. Liu, Y. Wong, Y. Zhang, W. Nie, Y. Su, and M. Kankanhalli, “Dual-stream recurrent neural network for video captioning,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , 2018
2018
Later among the works it cites.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Exploring visual relationship for image captioning,” in European Conference on Computer Vision (ECCV) , 2018, pp. 684–699
2018
Later among the works it cites.
Z. Yu, J. Yu, C. Xiang, J. Fan, and D. Tao, “Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering,” IEEE Transactions on Neural Networks and Learning Systems , vol. 29, no. 12, pp. 5947–5959, 2018
2018
Later among the works it cites.
D.-K. Nguyen and T. Okatani, “Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 6087–6096, 2018
2018
Later among the works it cites.
Z. Tu, W. Xie, J. Dauwels, B. Li, and J. Yuan, “Semantic cues enhanced multi-modality multi-stream cnn for action recognition,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , 2018
2018
Later among the works it cites.
D. Tao, Y. Guo, B. Yu, J. Pang, and Z. Yu, “Deep multi-view feature learning for person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , vol. 28, no. 10, pp. 2657–2666, 2018
2018
Later among the works it cites.
W. Jiang, L. Ma, Y.-G. Jiang, W. Liu, and T. Zhang, “Recurrent fusion network for image captioning,” in European Conference on Computer Vision (ECCV) , 2018, pp. 499–515
2018
Later among the works it cites.