Fetching the paper…
Reading the bibliography…
Image captioning can be improved if the structure of the graphical representations can be formulated with conceptual positional binding.
P. Smolensky, “Tensor product variable binding and the representation of symbolic structures in connectionist systems,” Artificial intelligence , vol. 46, no. 1-2, pp. 159–216, 1990
1990
Earlier work this paper cites.
A. Karpathy, A. Joulin, and L. F. Fei-Fei, “Deep fragment embeddings for bidirectional image sentence mapping,” in Advances in neural information processing systems , 2014, pp. 1889–1897
2014
Earlier work this paper cites.
C. Wang, H. Yang, C. Bartz, and C. Meinel, “Image captioning with deep bidirectional lstms,” in Proceedings of the 2016 ACM on Multimedia Conference . ACM, 2016, pp. 988–997
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: Lessons learned from the 2015 mscoco image captioning challenge,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 4, pp. 652–663, 2017
2017
Earlier work this paper cites.
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng, “Semantic compositional networks for visual captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , vol. 2, 2017
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , vol. 6, 2017, p. 2
2017
Cited alongside, same era.
Q. Wu, C. Shen, P. Wang, A. Dick, and A. van den Hengel, “Image captioning and visual question answering based on attributes and external knowledge,” IEEE transactions on pattern analysis and machine intelligence , 2017
2017
Cited alongside, same era.
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei, “Boosting image captioning with attributes,” in IEEE International Conference on Computer Vision, ICCV , 2017, pp. 22–29
2017
Cited alongside, same era.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Neural baby talk,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 7219–7228
2018
Later among the works it cites.
H. Chen, G. Ding, Z. Lin, S. Zhao, and J. Han, “Show, observe and tell: Attribute-driven attention model for image captioning.” in IJCAI , 2018, pp. 606–612
2018
Later among the works it cites.
Y. Li, W. Ouyang, B. Zhou, J. Shi, C. Zhang, and X. Wang, “Factorizable net: An efficient subgraph-based framework for scene graph generation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 335–351
2018
Later among the works it cites.
Q. Huang, P. Smolensky, X. He, L. Deng, and D. Wu, “Tensor product generation networks for deep nlp modeling,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , vol. 1, 2018, pp. 1263–1273
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in CVPR , vol. 1, no. 2, 2017, p. 3
2017
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR , vol. 3, no. 5, 2018, p. 6
2018
Cited alongside, same era.
C. Sur, P. Liu, Y. Zhou, and D. Wu, “Semantic tensor product for image captioning,” in Proceedings of the International Conference on Big Data Computing and Communication , 2019
2019
Closest in time.
H. Q, D. L, W. D, L. C, and H. X, “Attentive tensor product learning,” in Proceedings of The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI) , 2019
2019
Closest in time.