Fetching the paper…
Reading the bibliography…
The correlation between NLG automatic evaluation metrics and human evaluation is often regarded as a critical criterion for assessing the capability of an evaluation metric.
The proof and measurement of association between two things
C. Spearman. 1904 · 1904
Earlier work this paper cites.
The treatment of ties in ranking problems
M. G. KENDALL. 1945 · 1945
Earlier work this paper cites.
Computer intensive methods for hypothesis testing: An introduction
Eric W Noreen. 1989 · 1989
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Assessing the effect of inconsistent assessors on summarization evaluation
Karolina Owczarzak, Peter A. Rankel, Hoa Trang Dang, and John M. Conroy. 2012 · 2012
Earlier work this paper cites.
Metrics, statistics, tests
Tetsuya Sakai. 2013 · 2013
Earlier work this paper cites.
chrf: character n-gram f-score for automatic MT evaluation
Maja Popovic. 2015 · 2015
Earlier work this paper cites.
On the discriminative power of hyper-parameters in cross-validation and how to choose them
Vito Walter Anelli, Tommaso Di Noia, Eugenio Di Sciascio, Claudio Pomo, and Azzurra Ragone. 2019 · 2019
Earlier work this paper cites.
Revisiting online personal search metrics with the user in mind
Azin Ashkan and Donald Metzler. 2019 · 2019
Earlier work this paper cites.
Decline of pearson’sr with categorization of variables: a large-scale simulation
Takahiro Onoshima, Kenpei Shiina, Takashi Ueda, and Saori Kubo. 2019 · 2019
Earlier work this paper cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
Proceedings of the 3rd International Workshop on Natural Language Generation from the Semantic Web (WebNLG+) . Association for Computational Linguistics, Dublin, Ireland (Virtual)
Thiago Castro Ferreira, Claire Gardent, Nikolai Ilinykh, Chris van der Lee, Simon Mille, Diego Moussallem, and Anastasia Shimorina, editors. 2020 · 2020
Earlier work this paper cites.
Tangled up in BLEU: reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Earlier work this paper cites.
USR: an unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskénazi. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C. Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
BLEURT: learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh. 2020 · 2020
Cited alongside, same era.
Assessing ranking metrics in top-n recommendation
Daniel Valcarce, Alejandro Bellogín, Javier Parapar, and Pablo Castells. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
A statistical analysis of summarization evaluation metrics using resampling methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Results of WMT22 metrics shared task: Stop using BLEU - neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George F. Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Can large language models be an alternative to human evaluations?
David Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Later among the works it cites.
Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration
Daniel Deutsch, George F. Foster, and Markus Freitag. 2023 · 2023
Later among the works it cites.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frédéric Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George F. Foster. 2023 · 2023
Later among the works it cites.
Tigerscore: Towards building explainable metric for all text generation tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2021 · 2021
Cited alongside, same era.
Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George F. Foster, Alon Lavie, and Ondrej Bojar. 2021b · 2021
Cited alongside, same era.
Openmeva: A benchmark for evaluating open-ended story generation metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021 · 2021
Cited alongside, same era.
Evaluating evaluation measures for ordinal classification and ordinal quantification
Tetsuya Sakai. 2021a · 2021
Cited alongside, same era.
On the instability of diminishing return IR measures
Tetsuya Sakai. 2021b · 2021
Cited alongside, same era.
The statistical advantage of automatic NLG metrics at the system level
Johnny Tian-Zheng Wei and Robin Jia. 2021 · 2021
Cited alongside, same era.
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023 · 2023
Later among the works it cites.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023 · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Large language models are not yet human-level evaluators for abstractive summarization
Chenhui Shen, Liying Cheng, Xuan-Phi Nguyen, Yang You, and Lidong Bing. 2023 · 2023
Later among the works it cites.
Automated evaluation of personalized text generation using large language models
Yaqing Wang, Jiepu Jiang, Mingyang Zhang, Cheng Li, Yi Liang, Qiaozhu Mei, and Michael Bendersky. 2023 · 2023
Later among the works it cites.
INSTRUCTSCORE: towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023 · 2023
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Themis: Towards flexible and interpretable NLG evaluation
Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin, and Xiaojun Wan. 2024 · 2024
Closest in time.
Guardians of the machine translation meta-evaluation: Sentinel metrics fall in!
Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Edoardo Barba, and Roberto Navigli. 2024 · 2024
Closest in time.