Fetching the paper…
Reading the bibliography…
State-of-The-Art (SoTA) image captioning models are often trained on the MicroSoft Common Objects in Context (MS-COCO) dataset, which contains human-annotated captions with an average length of approximately ten tokens.
Brown T, Mann B, Ryder N, et al (2020) Language models are few-shot learners. In: Advances in Neural Information Processing Systems, pp 1877–1901
1901
Earlier work this paper cites.
Cho J, Lei J, Tan H, et al (2021) Unifying vision-and-language tasks via text generation. In: International Conference on Machine Learning (ICML), PMLR, pp 1931–1942
1942
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, et al (2002) Bleu: a method for automatic evaluation of machine translation. In: Annual meeting of the Association for Computational Linguistics, pp 311–318
2002
Earlier work this paper cites.
Banerjee S, Lavie A (2005) Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pp 65–72
2005
Earlier work this paper cites.
Smith NA, Eisner J (2005) Contrastive estimation: Training log-linear models on unlabeled data. In: Annual Meeting of the Association for Computational Linguistics (ACL), pp 354–362
2005
Earlier work this paper cites.
Ordonez V, Kulkarni G, Berg T (2011) Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems 24
2011
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, et al (2014) Microsoft coco: Common objects in context. In: European Conference on Computer Vision (ECCV), Springer, pp 740–755
2014
Earlier work this paper cites.
Young P, Lai A, Hodosh M, et al (2014) From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the association for computational linguistics 2:67–78
2014
Earlier work this paper cites.
Donahue J, Anne Hendricks L, Guadarrama S, et al (2015) Long-term recurrent convolutional networks for visual recognition and description. In: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp 2625–2634
2015
Earlier work this paper cites.
Karpathy A, Fei-Fei L (2015) Deep visual-semantic alignments for generating image descriptions. In: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp 3128–3137
2015
Earlier work this paper cites.
Vedantam R, Lawrence Zitnick C, Parikh D (2015) Cider: Consensus-based image description evaluation. In: Computer Vision and Pattern Recognition (CVPR), IEEE, pp 4566–4575
2015
Earlier work this paper cites.
Anderson P, Fernando B, Johnson M, et al (2016) Spice: Semantic propositional image caption evaluation. In: European Conference on Computer Vision (ECCV), Springer, pp 382–398
2016
Earlier work this paper cites.
Sun C, Shrivastava A, Singh S, et al (2017) Revisiting unreasonable effectiveness of data in deep learning era. In: International Conference on Computer Vision. IEEE, pp 843–852
2017
Earlier work this paper cites.
Anderson P, He X, Buehler C, et al (2018) Bottom-up and top-down attention for image captioning and visual question answering. In: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp 6077–6086
2018
Earlier work this paper cites.
Rohrbach A, Hendricks LA, Burns K, et al (2018) Object hallucination in image captioning. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp 4035–4045
2018
Earlier work this paper cites.
Sharma P, Ding N, Goodman S, et al (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 2556–2565
2018
Earlier work this paper cites.
Aneja J, Agrawal H, Batra D, et al (2019) Sequential latent spaces for modeling the intention during diverse image captioning. In: International Conference on Computer Vision (ICCV), IEEE/CVF, pp 4261–4270
2019
Cited alongside, same era.
Huang L, Wang W, Chen J, et al (2019) Attention on attention for image captioning. In: International Conference on Computer Vision (ICCV), IEEE/CVF, pp 4634–4643
2019
Cited alongside, same era.
Li X, Yin X, Li C, et al (2020) Oscar: Object-semantics aligned pre-training for vision-language tasks. In: European Conference on Computer Vision (ECCV), Springer, pp 121–137
2020
Cited alongside, same era.
Zhou C, Gu J, Neubig G (2020) Understanding knowledge distillation in non-autoregressive machine translation. In: International Conference on Learning Representations (ICLR)
2020
Cited alongside, same era.
Zha D, Bhat ZP, Lai KH, et al (2023) Data-centric ai: Perspectives and challenges. In: International Conference on Data Mining (SDM), SIAM, pp 945–948
2023
Closest in time.
Dong H, Li J, Wu B, et al (2024) Benchmarking and improving detail image caption. arXiv preprint arXiv:240519092
2024
Closest in time.
Fan L, Krishnan D, Isola P, et al (2024) Improving clip training with language rewrites. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Ge Y, Zeng X, Huffman JS, et al (2024) Visual fact checker: Enabling high-fidelity detailed caption generation. In: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE/CVF, pp 14033–14042
2024
Closest in time.
Li J, Vo DM, Sugimoto A, et al (2024) Evcap: Retrieval-augmented image captioning with external visual-name memory for open-world comprehension. In: Conference on Computer Vision and Pattern Recognition, IEEE/CVF, pp 13733–13742
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Changpinyo S, Sharma P, Ding N, et al (2021) Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In: Conference on Computer Vision and Pattern Recognition, IEEE/CVF, pp 3558–3568
2021
Cited alongside, same era.
Radford A, Kim JW, Hallacy C, et al (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML), PMLR, pp 8748–8763
2021
Cited alongside, same era.
Shi Z, Liu H, Zhu X (2021) Enhancing descriptive image captioning with natural language inference. In: Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pp 269–277
2021
Cited alongside, same era.
Wang S, Yao Z, Wang R, et al (2021) Faier: Fidelity and adequacy ensured image caption evaluation. In: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE/CVF, pp 14050–14059
2021
Cited alongside, same era.
Alayrac JB, Donahue J, Luc P, et al (2022) Flamingo: a visual language model for few-shot learning. In: Advances in Neural Information Processing Systems, pp 23716–23736
2022
Cited alongside, same era.
Chen Q, Deng C, Wu Q (2022) Learning distinct and representative modes for image captioning. In: Advances in Neural Information Processing Systems
2022
Cited alongside, same era.
Li J, Li D, Xiong C, et al (2022) BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning (ICML), PMLR
2022
Cited alongside, same era.
NLP Connect (2022) vit-gpt2-image-captioning (revision 0e334c7). 10.57967/hf/0222 , URL https://huggingface.co/nlpconnect/vit-gpt2-image-captioning
2022
Cited alongside, same era.
2024
Closest in time.
Luu DT, Le VT, Vo DM (2024) Questioning, answering, and captioning for zero-shot detailed image caption. In: Asian Conference on Computer Vision (ACCV), pp 242–259
2024
Closest in time.
Petryk S, Chan DM, Kachinthaya A, et al (2024) Aloha: A new measure for hallucination in captioning models. In: Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics
2024
Closest in time.
Rotstein N, Bensaïd D, Brody S, et al (2024) Fusecap: Leveraging large language models for enriched fused image captions. In: Winter Conference on Applications of Computer Vision (WACV), IEEE/CVF, pp 5677–5688
2024
Closest in time.
Wada Y, Kaneda K, Saito D, et al (2024) Polos: Multimodal metric learning from human feedback for image captioning. In: Conference on Computer Vision and Pattern Recognition, IEEE/CVF, pp 13559–13568
2024
Closest in time.
Yang X, Yang Y, Wu J, et al (2024) CA-captioner: A novel concentrated attention for image captioning. Expert Systems with Applications 250:123847
2024
Closest in time.
Yu Q, Sun Q, Zhang X, et al (2024) Capsfusion: Rethinking image-text data at scale. In: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE/CVF, pp 14022–14032
2024
Closest in time.
Zhu D, Chen J, Haydarov K, et al (2024) Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions. Transactions on Machine Learning Research
2024
Closest in time.
Gao N, Yao R, Chen P, et al (2025) Multi-granularity semantic relational mapping for image caption. Expert Systems with Applications 264:125847
2025
Closest in time.
Kim T, Lee S, Kim SW, et al (2025) Vipcap: Retrieval text-based visual prompts for lightweight image captioning. In: AAAI Conference on Artificial Intelligence, pp 4320–4328
2025
Closest in time.
Lai Z, Zhang H, Zhang B, et al (2025) Veclip: Improving clip training via visual-enriched captions. In: European Conference on Computer Vision, Springer, pp 111–127
2025
Closest in time.