Fetching the paper…
Reading the bibliography…
The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: BLEU: a Method for Automatic Evaluation of Machine Translation. In: ACL (2002)
2002
Earlier work this paper cites.
Lin, C.Y.: ROUGE: A Package for Automatic Evaluation of Summaries. In: ACL Workshops (2004)
2004
Earlier work this paper cites.
Banerjee, S., Lavie, A.: METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In: ACL Workshops (2005)
2005
Earlier work this paper cites.
Socher, R., Fei-Fei, L.: Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora. In: CVPR (2010)
2010
Earlier work this paper cites.
Yao, B.Z., Yang, X., Lin, L., Lee, M.W., Zhu, S.C.: I2t: Image parsing to text description. Proceedings of the IEEE 98
2010
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common Objects in context. In: ECCV (2014)
2014
Earlier work this paper cites.
Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: CVPR (2015)
2015
Earlier work this paper cites.
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: CIDEr: Consensus-Based Image Description Evaluation. In: CVPR (2015)
2015
Earlier work this paper cites.
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator. In: CVPR (2015)
2015
Earlier work this paper cites.
Anderson, P., Fernando, B., Johnson, M., Gould, S.: SPICE: Semantic Propositional Image Caption Evaluation. In: ECCV (2016)
2016
Earlier work this paper cites.
Rennie, S.J., Marcheret, E., Mroueh, Y., Ross, J., Goel, V.: Self-Critical Sequence Training for Image Captioning. In: CVPR (2017)
2017
Earlier work this paper cites.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: CVPR (2018)
2018
Earlier work this paper cites.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In: ACL (2018)
2018
Earlier work this paper cites.
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., Anderson, P.: nocaps: Novel object captioning at scale. In: ICCV (2019)
2019
Earlier work this paper cites.
Huang, L., Wang, W., Chen, J., Wei, X.Y.: Attention on Attention for Image Captioning. In: ICCV (2019)
2019
Earlier work this paper cites.
Yang, X., Tang, K., Zhang, H., Cai, J.: Auto-Encoding Scene Graphs for Image Captioning. In: CVPR (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. In: NeurIPS (2020)
2020
Earlier work this paper cites.
Cornia, M., Baraldi, L., Cucchiara, R.: SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability. In: ICRA (2020)
2020
Earlier work this paper cites.
Gurari, D., Zhao, Y., Zhang, M., Bhattacharya, N.: Captioning Images Taken by People Who Are Blind. In: ECCV (2020)
2020
Earlier work this paper cites.
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., Ferrari, V.: The Open Images Dataset V4. IJCV 128
2020
Earlier work this paper cites.
Sidorov, O., Hu, R., Rohrbach, M., Singh, A.: TextCaps: A Dataset for Image Captioning with Reading Comprehension. In: ECCV (2020)
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Hessel, J., Holtzman, A., Forbes, M., Bras, R.L., Choi, Y.: CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In: EMNLP (2021)
2021
Earlier work this paper cites.
Hu, E.J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.: LoRA: Low-Rank Adaptation of Large Language Models. In: ICLR (2021)
2021
Earlier work this paper cites.
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., Carreira, J.: Perceiver: General perception with iterative attention. In: ICML (2021)
2021
Cited alongside, same era.
Lester, B., Al-Rfou, R., Constant, N.: The Power of Scale for Parameter-Efficient Prompt Tuning. In: EMNLP (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning Transferable Visual Models From Natural Language Supervision. In: ICML (2021)
2023
Later among the works it cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual Instruction Tuning. In: NeurIPS (2023)
2023
Later among the works it cites.
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., Tang, J.: GPT understands, too. AI Open (2023)
2023
Later among the works it cites.
Ramos, R., Martins, B., Elliott, D., Kementchedjhieva, Y.: SmallCap: Lightweight Image Captioning Prompted With Retrieval Augmentation. In: CVPR (2023)
2023
Later among the works it cites.
Sarto, S., Barraco, M., Cornia, M., Baraldi, L., Cucchiara, R.: Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation. In: CVPR (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a Visual Language Model for Few-Shot Learning. In: NeurIPS (2022)
2022
Cited alongside, same era.
Barraco, M., Cornia, M., Cascianelli, S., Baraldi, L., Cucchiara, R.: The Unreasonable Effectiveness of CLIP Features for Image Captioning: An Experimental Analysis. In: CVPR Workshops (2022)
2022
Cited alongside, same era.
Cho, J., Yoon, S., Kale, A., Dernoncourt, F., Bui, T., Bansal, M.: Fine-grained Image Captioning with CLIP Reward. In: NAACL (2022)
2022
Cited alongside, same era.
Cornia, M., Baraldi, L., Cucchiara, R.: Explaining Transformer-based Image Captioning Models: An Empirical Analysis. AI Communications 35
2022
Cited alongside, same era.
Li, Y., Pan, Y., Yao, T., Mei, T.: Comprehending and Ordering Semantics for Image Captioning. In: CVPR (2022)
2022
Cited alongside, same era.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. In: NeurIPS (2022)
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Caffagni, D., Cocchi, F., Barsellotti, L., Moratelli, N., Sarto, S., Baraldi, L., Baraldi, L., Cornia, M., Cucchiara, R.: The Revolution of Multimodal Large Language Models: A Survey. In: ACL Findings (2024)
2024
Closest in time.
Caffagni, D., Cocchi, F., Moratelli, N., Sarto, S., Cornia, M., Baraldi, L., Cucchiara, R.: Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs. In: CVPR Workshops (2024)
2024
Closest in time.
Cha, J., Kang, W., Mun, J., Roh, B.: Honeybee: Locality-enhanced Projector for Multimodal LLM. In: CVPR (2024)
2024
Closest in time.
Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., Wang, L.: Aligning Large Multi-Modal Model with Robust Instruction Tuning. In: ICLR (2024)
2024
Closest in time.
Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved Baselines with Visual Instruction Tuning. In: CVPR (2024)
2024
Closest in time.
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., Lee, Y.J.: LLaVA-NeXT: Improved reasoning, OCR, and world knowledge (2024)
2024
Closest in time.
Liu, S.Y., Wang, C.Y., Yin, H., Molchanov, P., Wang, Y.C.F., Cheng, K.T., Chen, M.H.: DoRA: Weight-Decomposed Low-Rank Adaptation. In: ICML (2024)
2024
Closest in time.
Moratelli, N., Barraco, M., Cornia, M., Baraldi, L., Cucchiara, R.: Are Learnable Prompts the Right Way of Prompting? Adapting Vision-and-Language Models with Memory Optimization. IEEE Intelligent Systems (2024)
2024
Closest in time.
Moratelli, N., Caffagni, D., Cornia, M., Baraldi, L., Cucchiara, R.: Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization. In: BMVC (2024)
2024
Closest in time.
Sarto, S., Cornia, M., Baraldi, L., Cucchiara, R.: BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues. In: ECCV (2024)
2024
Closest in time.
2024
Closest in time.
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., Xie, S.: Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. In: CVPR (2024)
2024
Closest in time.
Wang, B., Wu, F., Han, X., Peng, J., Zhong, H., Zhang, P., Dong, X., Li, W., Li, W., Wang, J., et al.: VIGC: Visual Instruction Generation and Correction. In: AAAI (2024)
2024
Closest in time.