Fetching the paper…
Reading the bibliography…
Most research about natural language generation (NLG) relies on evaluation benchmarks with limited references for a sample, which may result in poor correlations with human judgements.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Paraphrasing for automatic evaluation
David Kauchak and Regina Barzilay. 2006 · 2006
Earlier work this paper cites.
Effective academic writing: the short essay
Alice Savage and Patricia Mayer. 2006 · 2006
Earlier work this paper cites.
Re-evaluating machine translation results with paraphrase support
Liang Zhou, Chin-Yew Lin, and Eduard Hovy. 2006a · 2006
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
BLEU is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Earlier work this paper cites.
Multi-reference training with pseudo-references for neural translation and text generation
Renjie Zheng, Mingbo Ma, and Liang Huang. 2018 · 2018
Earlier work this paper cites.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey Bigham. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
A study in improving BLEU reference coverage with diverse automatic paraphrasing
Rachel Bawden, Biao Zhang, Lisa Yankovskaya, Andre Tättar, and Matt Post. 2020b · 2020
Earlier work this paper cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020b · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Simulated multiple reference training improves low-resource machine translation
Huda Khayrallah, Brian Thompson, Matt Post, and Philipp Koehn. 2020 · 2020
Cited alongside, same era.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020a · 2020
Findings of the 2022 conference on machine translation (WMT22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022 · 2022
Later among the works it cites.
A survey of pretrained language models based text generation
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2022 · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022 · 2022
Later among the works it cites.
A survey of evaluation metrics used for nlg systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020b · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger. 2020 · 2020
Cited alongside, same era.
The (un)suitability of automatic evaluation metrics for text simplification
Fernando Alva-Manchego, Carolina Scarton, and Lucia Specia. 2021 · 2021
Cited alongside, same era.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023 · 2023
Closest in time.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Closest in time.
Ties matter: Modifying kendall’s tau for modern metric meta-evaluation
Daniel Deutsch, George Foster, and Markus Freitag. 2023 · 2023
Closest in time.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Closest in time.
Is chatgpt a good translator? yes with gpt-4 as the engine
WX Jiao, WX Wang, JT Huang, Xing Wang, and ZP Tu. 2023 · 2023
Closest in time.
Reducing sequence length by predicting edit operations with large language models
Masahiro Kaneko and Naoaki Okazaki. 2023 · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
Qingyu Lu, Baopu Qiu, Liang Ding, Liping Xie, and Dacheng Tao. 2023 · 2023
Closest in time.
Chatgpt as a factual inconsistency evaluator for abstractive text summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023 · 2023
Closest in time.
Is chatgpt a good nlg evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023 · 2023
Closest in time.
Benchmarking large language models for news summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto. 2023 · 2023
Closest in time.