Fetching the paper…
Reading the bibliography…
We address a fundamental challenge in Natural Language Generation (NLG) model evaluation -- the design and evaluation of evaluation metrics.
General intelligence, objectively determined and measured
C Spearman. 1904 · 1904
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 1909
Earlier work this paper cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 1909
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Correlation calculated from faulty data
Charles Spearman. 1910 · 1910
Earlier work this paper cites.
Multiple-factor analysis; a development and expansion of the vectors of mind
Louis Leon Thurstone. 1947 · 1947
Earlier work this paper cites.
Coefficient alpha and the internal structure of tests
Lee J Cronbach. 1951 · 1951
Earlier work this paper cites.
Construct validity in psychological tests
Lee J Cronbach and Paul E Meehl. 1955 · 1955
Earlier work this paper cites.
Convergent and discriminant validation by the multitrait-multimethod matrix
Donald T Campbell and Donald W Fiske. 1959 · 1959
Earlier work this paper cites.
A general structural equation model with dichotomous, ordered categorical, and continuous latent variable indicators
Bengt Muthén. 1984 · 1984
Earlier work this paper cites.
Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning
Samuel Messick. 1995 · 1995
Earlier work this paper cites.
Introduction to measurement theory
Mary J Allen and Wendy M Yen. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Fill in the blanc: Human-free quality estimation of document summaries
Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Supert: Towards new frontiers in unsupervised evaluation metrics for multi-document summarization
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Cited alongside, same era.
Paraphrasing for automatic evaluation
David Kauchak and Regina Barzilay. 2006 · 2006
Cited alongside, same era.
Defining and evaluating fair natural language generation
Catherine Yeo and Alyssa Chen. 2020 · 2008
Cited alongside, same era.
Curious case of language generation evaluation metrics: A cautionary tale
Ozan Caglayan, Pranava Madhyastha, and Lucia Specia. 2020 · 2010
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Cited alongside, same era.
The reprogen shared task on reproducibility of human evaluations in nlg: Overview and results
Anja Belz, Anastasia Shimorina, Shubham Agarwal, and Ehud Reiter. 2021 · 2021
Later among the works it cites.
All that’s ‘human’is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Later among the works it cites.
A statistical analysis of summarization evaluation metrics using resampling methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
A fine-grained analysis of bertscore
Michael Hanna and Ondřej Bojar. 2021 · 2021
Later among the works it cites.
Genie: Toward reproducible and standardized human evaluation for text generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015 · 2015
Cited alongside, same era.
Better summarization evaluation with word embeddings for rouge
Jun-Ping Ng and Viktoria Abrecht. 2015 · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Chia-Wei Liu, Ryan Lowe, Iulian V Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
chrf++: words helping character n-grams
Maja Popović. 2017 · 2017
Cited alongside, same era.
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld. 2021 · 2021
Later among the works it cites.
Perturbation checklists for evaluating nlg evaluation metrics
Ananya B Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M Khapra. 2021 · 2021
Later among the works it cites.
Societal biases in language generation: Progress and challenges
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Later among the works it cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Yixin Liu, Alexander R Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, et al. 2022 · 2022
Later among the works it cites.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2022 · 2022
Later among the works it cites.
On the effectiveness of automated metrics for text generation systems
Pius Von Däniken, Jan Deriu, Don Tuggener, and Mark Cieliebak. 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexandra Olteanu. 2022 · 2022
Later among the works it cites.
Seahorse: A multilingual, multifaceted dataset for summarization evaluation
Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, and Ankur P Parikh. 2023 · 2023
Closest in time.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
Nlg evaluation metrics beyond correlation analysis: An empirical metric preference checklist
Iftitahu Ni’mah, Meng Fang, Vlado Menkovski, and Mykola Pechenizkiy. 2023 · 2023
Closest in time.