Fetching the paper…
Reading the bibliography…
To prevent the costly and inefficient use of resources on low-quality annotations, we want a method for creating a pool of dependable annotators who can effectively complete difficult tasks, such as evaluating automatic summarization.
The proof and measurement of association between two things
C. Spearman. 1904 · 1904
Earlier work this paper cites.
A technique for the measurement of attitudes
Rensis Likert. 1932 · 1932
Earlier work this paper cites.
Design of experiments
Ronald Aylmer Fisher. 1936 · 1936
Earlier work this paper cites.
Significance tests which may be applied to samples from any populations
E. J. G. Pitman. 1937 · 1937
Earlier work this paper cites.
Nonparametric estimation from incomplete observations
E. L. Kaplan and Paul Meier. 1958 · 1958
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Bradley Efron. 1992 · 1992
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Answering the call for a standard reliability measure for coding data
Andrew F Hayes and Klaus Krippendorff. 2007 · 2007
Earlier work this paper cites.
Fast, cheap, and creative: Evaluating translation quality using Amazon’s Mechanical Turk
Chris Callison-Burch. 2009 · 2009
Earlier work this paper cites.
Instructional manipulation checks: Detecting satisficing to increase statistical power
Daniel M. Oppenheimer, Tom Meyvis, and Nicolas Davidenko. 2009 · 2009
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu. 2010 · 2010
Earlier work this paper cites.
Automatic evaluation of translation quality for distant language pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, and Hajime Tsukada. 2010 · 2010
Earlier work this paper cites.
Running experiments on amazon mechanical turk
Gabriele Paolacci, Jesse Chandler, and Panagiotis G Ipeirotis. 2010 · 2010
Cited alongside, same era.
Amazon’s mechanical turk: A new source of inexpensive, yet high-quality data?
Michael Buhrmester, Tracy Kwang, and Samuel D Gosling. 2011 · 2011
Cited alongside, same era.
Evaluating online labor markets for experimental research: Amazon.com’s mechanical turk
Adam J. Berinsky, Gregory A. Huber, and Gabriel S. Lenz. 2012 · 2012
Cited alongside, same era.
Learning whom to trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013 · 2013
Cited alongside, same era.
Machine reading tea leaves: Automatically evaluating topic coherence and topic model quality
Jey Han Lau, David Newman, and Timothy Baldwin. 2014 · 2014
Cited alongside, same era.
Turking overtime: How participant characteristics and behavior vary over time and day on amazon mechanical turk
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
A human evaluation of AMR-to-English generation systems
Emma Manning, Shira Wein, and Nathan Schneider. 2020 · 2020
Later among the works it cites.
Creativity on paid crowdsourcing platforms
Jonas Oppenlaender, Kristy Milland, Aku Visuri, Panos Ipeirotis, and Simo Hosio. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Antonio A Arechar, Gordon T Kraft-Todd, and David G Rand. 2017 · 2017
Cited alongside, same era.
Can machine translation systems be evaluated by the crowd alone
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2017 · 2017
Cited alongside, same era.
Evaluation of automatic video captioning using direct assessment
Yvette Graham, George Awad, and Alan Smeaton. 2018 · 2018
Cited alongside, same era.
Comparing Bayesian models of annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
HighRES: Highlight-based reference-less evaluation of summarization
Hardy Hardy, Shashi Narayan, and Andreas Vlachos. 2019 · 2019
Cited alongside, same era.
The second multilingual surface realisation shared task (SR’19): Overview and evaluation results
Simon Mille, Anja Belz, Bernd Bohnet, Yvette Graham, and Leo Wanner. 2019 · 2019
Cited alongside, same era.
Jessica Huynh, Jeffrey Bigham, and Maxine Eskenazi. 2021 · 2021
Later among the works it cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Quantifying and avoiding unfair qualification labour in crowdsourcing
Jonathan K. Kummerfeld. 2021 · 2021
Later among the works it cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Closest in time.
Shikib Mehri, Jinho Choi, L. F. D’Haro, Jan Deriu, Maxine Eskénazi, Milica Gasic, Kallirroi Georgila, Dilek Z. Hakkani-Tür, Zekang Li, Verena Rieser, Samira Shaikh, David R. Traum, Yi-Ting Yeh, Zhou Yu, Yizhe Zhang, and Chen Zhang. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Closest in time.
The human evaluation datasheet: A template for recording details of human evaluation experiments in NLP
Anastasia Shimorina and Anya Belz. 2022 · 2022
Closest in time.
Too good to be true: Bots and bad data from mechanical turk
Margaret A Webb and June P Tangney. 2022 · 2022
Closest in time.
OpenAI. 2023 · 2023
Closest in time.