Fetching the paper…
Reading the bibliography…
Human ratings are the gold standard in NLG evaluation.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
A law of comparative judgment
Louis L Thurstone. 1927 · 1927
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Theory of games and economic behavior
Oskar Morgenstern and John Von Neumann. 1953 · 1953
Earlier work this paper cites.
Magnitude estimation: notes on what, how, when, and why to use it
Howard R Moskowitz. 1977 · 1977
Earlier work this paper cites.
St. petersburg paradoxes: Defanged, dissected, and historically described
Paul A Samuelson. 1977 · 1977
Earlier work this paper cites.
The so-called allais paradox and rational decisions under uncertainty
Maurice Allais. 1979 · 1979
Earlier work this paper cites.
A simple sequentially rejective multiple test procedure
Sture Holm. 1979 · 1979
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
D Kahneman and A Tversky. 1979 · 1979
Earlier work this paper cites.
Recovering cardinal utility
Philip Dybvig and Heraklis Polemarchakis. 1981 · 1981
Earlier work this paper cites.
Magnitude estimation of linguistic acceptability
Ellen Gurman Bard, Dan Robertson, and Antonella Sorace. 1996 · 1996
Earlier work this paper cites.
Narrative prose generation
Charles B Callaway and James C Lester. 2002 · 2002
Earlier work this paper cites.
On the foundations of expected expected utility
Craig Boutilier. 2003 · 2003
Earlier work this paper cites.
Likert scales: How to (ab) use them?
Susan Jamieson. 2004 · 2004
Earlier work this paper cites.
Preference learning with gaussian processes
Wei Chu and Zoubin Ghahramani. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Cited alongside, same era.
TrueSkill™: a bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. 2006 · 2006
Cited alongside, same era.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Shaw Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Cited alongside, same era.
Introducing shared tasks to NLG: The TUNA shared task evaluation challenges
Albert Gatt and Anja Belz. 2009 · 2009
Cited alongside, same era.
Discrete vs. continuous rating scales for language evaluation in NLP
Anja Belz and Eric Kow. 2011 · 2011
Cited alongside, same era.
Analyzing and interpreting data from likert-type scales
Gail M Sullivan and Anthony R Artino Jr. 2013 · 2013
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Importance of search and evaluation strategies in neural dialogue modeling
Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019 · 2019
Later among the works it cites.
Some biases in likert scaling usage and its correction
J Pimentel and JL Pimentel. 2019 · 2019
Later among the works it cites.
Best practices for the human evaluation of automatically generated text
Chris Van Der Lee, Albert Gatt, Emiel Van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019 · 2019
Later among the works it cites.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient elicitation of annotations for human evaluation of machine translation
Keisuke Sakaguchi, Matt Post, and Benjamin Van Durme. 2014 · 2014
Cited alongside, same era.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, et al. 2016 · 2016
Cited alongside, same era.
Findings of the 2017 conference on machine translation WMT17
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, et al. 2017 · 2017
Cited alongside, same era.
Findings of the E2E NLG challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Cited alongside, same era.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
David M Howcroft, Anja Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A Hasan, Saad Mahamood, Simon Mille, Emiel Van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
Automatic detection of generated text is easiest when humans are fooled
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020 · 2020
Later among the works it cites.
Principles of economics
N Gregory Mankiw. 2020 · 2020
Later among the works it cites.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, et al. 2021 · 2021
Later among the works it cites.
Genie: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld. 2021 · 2021
Later among the works it cites.
Findings of the WMT 2021 shared task on quality estimation
Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André Martins. 2021 · 2021
Later among the works it cites.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Later among the works it cites.
Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexandra Olteanu. 2022 · 2022
Closest in time.