Fetching the paper…
Reading the bibliography…
Evaluation metric of visual captioning is important yet not thoroughly explored.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, M.; Young, P.; and Hockenmaier, J. 2013 · 2013
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
SPICE: Semantic propositional image caption evaluation
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016 · 2016
Earlier work this paper cites.
Learning to evaluate image captioning
Cui, Y.; Yang, G.; Veit, A.; Huang, X.; and Belongie, S. 2018 · 2018
Earlier work this paper cites.
Learning to evaluate image captioning
Huang, X.; Veit, A.; Huang, Z. A.; and Belongie, S. 2019 · 2019
Cited alongside, same era.
TIGEr: Text-to-Image Grounded Evaluator for Image Captioning
Jiang, Z.; Gan, C.; Wu, J.; Zhao, H.; and Xie, L. 2019 · 2019
Cited alongside, same era.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Zhao, W.; Peyrard, M.; Liu, F.; Gao, Y.; Meyer, C. M.; and Eger, S. 2019 · 2019
Cited alongside, same era.
ViLBERTScore: Evaluating Image Caption Using Vision-and-Language BERT
Lee, H.; Yoon, S.; Dernoncourt, F.; Kim, D. S.; Bui, T.; and Jung, K. 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; et al. 2020 · 2020
Cited alongside, same era.
Revisiting BERTScore for Image Captioning: Extensions and Relevance
EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching
Shi, Y.; Yang, X.; Xu, H.; Yuan, C.; Li, B.; Hu, W.; and Zha, Z. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Later among the works it cites.
CLAIR: Evaluating Image Captions with Large Language Models
Chan, D. M.; Petryk, S.; Gonzalez, J. E.; Darrell, T.; and Canny, J. 2023 · 2023
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2020
Cited alongside, same era.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
UMIC: Unreferenced Metric for Image Captioning
Lee, C. Y.; Yoon, J.; Dernoncourt, F.; Bui, T.; and Jung, K. 2021 · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Cited alongside, same era.
Positive-augmented contrastive learning for image and video captioning evaluation
Sarto, S.; Barraco, M.; Cornia, M.; Baraldi, L.; and Cucchiara, R. 2023 · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Later among the works it cites.
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
Lee, Y.; Park, I.; and Kang, M. 2024 · 2024
Closest in time.
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
Wada, Y.; Kaneda, K.; Saito, D.; and Sugiura, K. 2024 · 2024
Closest in time.