Fetching the paper…
Reading the bibliography…
Evaluating the compatibility between textual descriptions and corresponding images represents a core endeavor within multi-modal research.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: an automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics (extended abstract)
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
SPICE: semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017 · 2017
Earlier work this paper cites.
Shedding the cobra effect: problematising thematic emergence, triangulation, saturation and member checking
Lara Varpio, Rola Ajjawi, Lynn V Monrouxe, Bridget C O’Brien, and Charlotte E Rees. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Tiger: Text-to-image grounding for image caption evaluation
Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, and Jianfeng Gao. 2019 · 2019
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022 · 2022
Later among the works it cites.
Probing cross-modal semantics alignment capability from the textual perspective
Zheng Ma, Shi Zong, Mianzhi Pan, Jianbing Zhang, Shujian Huang, Xinyu Dai, and Jiajun Chen. 2022 · 2022
Later among the works it cites.
OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022 · 2022
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. 2022 · 2022
Later among the works it cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge aware semantic concept expansion for image-text matching
Botian Shi, Lei Ji, Pan Lu, Zhendong Niu, and Nan Duan. 2019 · 2019
Cited alongside, same era.
UNITER: universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021 · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
UMIC: an unreferenced metric for image captioning via contrastive learning
Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Trung Bui, and Kyomin Jung. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Fine-grained image captioning with CLIP reward
Jaemin Cho, Seunghyun Yoon, Ajinkya Kale, Franck Dernoncourt, Trung Bui, and Mohit Bansal. 2022 · 2022
Cited alongside, same era.
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Sharegpt4v: Improving large multi-modal models with better captions
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023 · 2023
Later among the works it cites.
Beyond generic: Enhancing image captioning with real-world knowledge using vision-language pre-training model
Kanzhi Cheng, Wenpo Song, Zheng Ma, Wenhao Zhu, Zixuan Zhu, and Jianbing Zhang. 2023 · 2023
Later among the works it cites.
Infometic: An informative metric for reference-free image caption evaluation
Anwen Hu, Shizhe Chen, Liang Zhang, and Qin Jin. 2023 · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Later among the works it cites.
Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation
Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. 2023 · 2023
Later among the works it cites.
Bounding and filling: A fast and flexible framework for image captioning
Zheng Ma, Changxin Wang, Bo Huang, Zixuan Zhu, and Jianbing Zhang. 2023 · 2023
Later among the works it cites.