Fetching the paper…
Reading the bibliography…
Image captioning models are usually trained according to human annotated ground-truth captions, which could generate accurate but generic captions.
Frankel, C., Swain, M.J., Athitsos, V.: Webseer: An image search engine for the world wide web. Tech. rep., Technical Report 96-14, University of Chicago, Computer Science Department (1996)
1996
Earlier work this paper cites.
Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural computation 9
1997
Earlier work this paper cites.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics. pp. 311–318 (2002)
2002
Earlier work this paper cites.
Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out. pp. 74–81 (2004)
2004
Earlier work this paper cites.
Banerjee, S., Lavie, A.: Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization. pp. 65–72 (2005)
2005
Earlier work this paper cites.
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Bengio, S., Vinyals, O., Jaitly, N., Shazeer, N.: Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems 28
2015
Earlier work this paper cites.
Cho, K., Courville, A., Bengio, Y.: Describing multimedia content using attention-based encoder-decoder networks. IEEE Transactions on Multimedia 17
2015
Earlier work this paper cites.
Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3128–3137 (2015)
2015
Earlier work this paper cites.
Ma, L., Lu, Z., Shang, L., Li, H.: Multimodal convolutional neural networks for matching image and sentence. In: Proceedings of the IEEE international conference on computer vision. pp. 2623–2631 (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28
2015
Cited alongside, same era.
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: Cider: Consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4566–4575 (2015)
2015
Cited alongside, same era.
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3156–3164 (2015)
2015
Cited alongside, same era.
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: International conference on machine learning. pp. 2048–2057. PMLR (2015)
2015
Cited alongside, same era.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 2556–2565 (2018)
2018
Later among the works it cites.
Yao, T., Pan, Y., Li, Y., Mei, T.: Exploring visual relationship for image captioning. In: Proceedings of the European conference on computer vision (ECCV). pp. 684–699 (2018)
2018
Later among the works it cites.
Liu, L., Tang, J., Wan, X., Guo, Z.: Generating diverse and descriptive image captions using visual paraphrases. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4240–4249 (2019)
2019
Later among the works it cites.
Makav, B., Kılıç, V.: A new image captioning approach for visually impaired people. In: 2019 11th International Conference on Electrical and Electronics Engineering (ELECO). pp. 945–949. IEEE (2019)
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anderson, P., Fernando, B., Johnson, M., Gould, S.: Spice: Semantic propositional image caption evaluation. In: European conference on computer vision. pp. 382–398. Springer (2016)
2016
Cited alongside, same era.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
2016
Cited alongside, same era.
Dai, B., Fidler, S., Urtasun, R., Lin, D.: Towards diverse and natural image descriptions via a conditional gan. In: Proceedings of the IEEE international conference on computer vision. pp. 2970–2979 (2017)
2017
Cited alongside, same era.
Dai, B., Lin, D.: Contrastive learning for image captioning. Advances in Neural Information Processing Systems 30
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Rennie, S.J., Marcheret, E., Mroueh, Y., Ross, J., Goel, V.: Self-critical sequence training for image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 7008–7024 (2017)
2017
Cited alongside, same era.
Shetty, R., Rohrbach, M., Anne Hendricks, L., Fritz, M., Schiele, B.: Speaking the same language: Matching machine to human captions by adversarial training. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 4135–4144 (2017)
2017
Cited alongside, same era.
Yao, T., Pan, Y., Li, Y., Qiu, Z., Mei, T.: Boosting image captioning with attributes. In: Proceedings of the IEEE international conference on computer vision. pp. 4894–4902 (2017)
2017
Cited alongside, same era.
Later among the works it cites.
Wang, Q., Chan, A.B.: Describing like humans: on diversity in image captioning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4195–4203 (2019)
2019
Later among the works it cites.
Xiong, Y., Du, B., Yan, P.: Reinforced transformer for medical image captioning. In: International Workshop on Machine Learning in Medical Imaging. pp. 673–680. Springer (2019)
2019
Later among the works it cites.
Cornia, M., Stefanini, M., Baraldi, L., Cucchiara, R.: Meshed-memory transformer for image captioning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10578–10587 (2020)
2020
Later among the works it cites.
Pan, Y., Yao, T., Li, Y., Mei, T.: X-linear attention networks for image captioning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10971–10980 (2020)
2020
Later among the works it cites.
Wang, J., Xu, W., Wang, Q., Chan, A.B.: Compare and reweight: Distinctive image captioning using similar images sets. In: European Conference on Computer Vision. pp. 370–386. Springer (2020)
2020
Later among the works it cites.
Wang, Z., Feng, B., Narasimhan, K., Russakovsky, O.: Towards unique and informative captioning of images. In: European Conference on Computer Vision. pp. 629–644. Springer (2020)
2020
Later among the works it cites.
Li, Y., Pan, Y., Chen, J., Yao, T., Mei, T.: X-modaler: A versatile and high-performance codebase for cross-modal analytics. In: Proceedings of the 29th ACM International Conference on Multimedia. pp. 3799–3802 (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning. pp. 8748–8763. PMLR (2021)
2021
Later among the works it cites.
Wang, J., Xu, W., Wang, Q., Chan, A.B.: Group-based distinctive image captioning with memory attention. In: Proceedings of the 29th ACM International Conference on Multimedia. pp. 5020–5028 (2021)
2021
Later among the works it cites.