Fetching the paper…
Reading the bibliography…
The ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans.
H. W. Kuhn, “The Hungarian method for the assignment problem,” Naval Research Logistics Quarterly , vol. 2, no. 1-2, pp. 83–97, 1955
1955
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proceedings of the Annual Meeting on Association for Computational Linguistics , 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Proceedings of the Annual Meeting on Association for Computational Linguistics Workshops , 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in Proceedings of the Annual Meeting on Association for Computational Linguistics Workshops , 2005
2005
Earlier work this paper cites.
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu, “I2t: Image parsing to text description,” in Proceedings of the IEEE , 2010
2010
Earlier work this paper cites.
R. Socher and L. Fei-Fei, “Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2010
2010
Earlier work this paper cites.
J. Pennington, R. Socher, and C. Manning, “GloVe: Global vectors for word representation,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2014
2014
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proceedings of the International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proceedings of the International Conference on Machine Learning , 2015
2015
Earlier work this paper cites.
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba, “Sequence level training with recurrent neural networks,” in Proceedings of the International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based Image Description Evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
J. Johnson, A. Karpathy, and L. Fei-Fei, “Densecap: Fully convolutional localization networks for dense captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016
2016
Cited alongside, same era.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2018
2018
Later among the works it cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Neural Baby Talk,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018
2018
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018
2018
Later among the works it cites.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Exploring Visual Relationship for Image Captioning,” in Proceedings of the European Conference on Computer Vision , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017
2017
Cited alongside, same era.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of spider,” in Proceedings of the International Conference on Computer Vision , 2017
2017
Cited alongside, same era.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 4, pp. 652–663, 2017
2017
Cited alongside, same era.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Cited alongside, same era.
M. Pedersoli, T. Lucas, C. Schmid, and J. Verbeek, “Areas of attention for image captioning,” in Proceedings of the International Conference on Computer Vision , 2017
2017
Cited alongside, same era.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei, “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations,” International Journal of Computer Vision , vol. 123, no. 1, pp. 32–73, 2017
2017
Cited alongside, same era.
2018
Cited alongside, same era.
W. Jiang, L. Ma, Y.-G. Jiang, W. Liu, and T. Zhang, “Recurrent Fusion Network for Image Captioning,” in Proceedings of the European Conference on Computer Vision , 2018
2018
Later among the works it cites.
M. Cornia, L. Baraldi, G. Serra, and R. Cucchiara, “Paying more attention to saliency: Image captioning with saliency and context attention,” ACM Transactions on Multimedia Computing, Communications, and Applications , vol. 14, no. 2, p. 48, 2018
2018
Later among the works it cites.
J. Aneja, A. Deshpande, and A. G. Schwing, “Convolutional image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
F. Landi, L. Baraldi, M. Corsini, and R. Cucchiara, “Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters,” in Proceedings of the British Machine Vision Conference , 2019
2019
Closest in time.
M. Cornia, L. Baraldi, and R. Cucchiara, “Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019
2019
Closest in time.
S. Herdade, A. Kappeler, K. Boakye, and J. Soares, “Image Captioning: Transforming Objects into Words,” in Advances in Neural Information Processing Systems , 2019
2019
Closest in time.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics , 2019
2019
Closest in time.
2019
Closest in time.
X. Yang, K. Tang, H. Zhang, and J. Cai, “Auto-Encoding Scene Graphs for Image Captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019
2019
Closest in time.