Fetching the paper…
Reading the bibliography…
The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation.
Rank correlation methods
Maurice George Kendall · 1948
Earlier work this paper cites.
Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm
A. P. Dawid and A. M. Skene · 1979
Earlier work this paper cites.
Inferring Ground Truth from Subjective Labelling of Venus Images
Padhraic Smyth, Usama Fayyad, Michael Burl, Pietro Perona, and Pierre Baldi · 1994
Earlier work this paper cites.
Implementing crowdsourcing-based relevance experimentation: an industrial perspective
Omar Alonso · 2013
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel · 2013
Earlier work this paper cites.
Learning whom to trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy · 2013
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri · 2016
Earlier work this paper cites.
Probabilistic modeling for crowdsourcing partially-subjective ratings
An Nguyen, Matthew Halpern, Byron Wallace, and Matthew Lease · 2016
Earlier work this paper cites.
The calibrated sigma method: An efficient remedy for between-group differences in response category use on likert scales
Bert Weijters, Hans Baumgartner, and Maggie Geuens · 2016
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
Efficient online scalar annotation with bounded support
Keisuke Sakaguchi and Benjamin Van Durme · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
ELI5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Earlier work this paper cites.
Faithful multimodal explanation for visual question answering
Jialin Wu and Raymond Mooney · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal · 2020
Earlier work this paper cites.
Inquisitive question generation for high level text comprehension
Wei-Jen Ko, Te-yuan Chen, Yiyan Huang, Greg Durrett, and Junyi Jessy Li · 2020
Earlier work this paper cites.
ALICE: Active learning with contrastive natural language explanations
Weixin Liang, James Zou, and Zhou Yu · 2020
Earlier work this paper cites.
ExpBERT: Representation engineering with natural language explanations
Shikhar Murty, Pang Wei Koh, and Percy Liang · 2020
Earlier work this paper cites.
WT5?! Training Text-to-Text Models to Explain their Predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan · 2020
Cited alongside, same era.
We need to consider disagreement in evaluation
Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, Alexandra Uma, et al · 2021
Cited alongside, same era.
Did they answer? subjective acts and intents in conversational discourse
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, and Katrin Erk · 2021
Cited alongside, same era.
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S. Bernstein · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu · 2023
Closest in time.
Can Large Language Models Be an Alternative to Human Evaluations?
Cheng-Han Chiang and Hung-yi Lee · 2023
Closest in time.
ChatGPT outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann · 2023
Closest in time.
Evaluating Verifiability in Generative Search Engines
Nelson F. Liu, Tianyi Zhang, and Percy Liang · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Semeval-2021 task 12: Learning with disagreements
Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller, Jon Chamberlain, Barbara Plank, Edwin Simpson, and Massimo Poesio · 2021
Cited alongside, same era.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasovic · 2021
Cited alongside, same era.
Content moderation as a political issue: The twitter discourse around trump’s ban
Meysam Alizadeh, Fabrizio Gilardi, Emma Hoes, K Jonathan Klüser, Mael Kubli, and Nahema Marchal · 2022
Cited alongside, same era.
Is GPT-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi · 2022
Cited alongside, same era.
The authenticity gap in human evaluation
Kawin Ethayarajh and Dan Jurafsky · 2022
Cited alongside, same era.
Your answer is incorrect… would you like to know why? introducing a bilingual short answer feedback dataset
Anna Filighera, Siddharth Parihar, Tim Steuer, Tobias Meuser, and Sebastian Ochs · 2022
Cited alongside, same era.
SNaC: Coherence error detection for narrative summarization
Tanya Goyal, Junyi Jessy Li, and Greg Durrett · 2022
Cited alongside, same era.
Petter Törnberg · 2023
Closest in time.
Is ChatGPT a good NLG evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou · 2023
Closest in time.
Automated metrics for medical multi-document summarization disagree with human evaluations
Lucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong, Bailey Kuehl, Erin Bransom, and Byron Wallace · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
A critical evaluation of evaluations for long-form question answering
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi · 2023
Closest in time.
Evaluating subjective cognitive appraisals of emotions from large language models
Hongli Zhan, Desmond Ong, and Junyi Jessy Li · 2023
Closest in time.
Benchmarking large language models for news summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B. Hashimoto · 2023
Closest in time.
Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson · 2023
Closest in time.
Complex claim verification with evidence retrieved in the wild
Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, and Eunsol Choi · 2024
Closest in time.
AnnoLLM: Making large language models to be better crowdsourced annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, A-Long Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, and Weizhu Chen · 2024
Closest in time.
TIGERScore: Towards building explainable metric for all text generation tasks
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen · 2024
Closest in time.
Evallm: Interactive evaluation of large language model prompts on user-defined criteria
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim · 2024
Closest in time.
SemEval-2017 task 4: Sentiment analysis in Twitter
Sara Rosenthal, Noura Farra, and Preslav Nakov · 2088
Closest in time.