Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks, where relying on human evaluators can be costly, time-consuming, and unscalable.
Content and style in personality assessment
Douglas N. Jackson and Samuel Messick. 1958 · 1958
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E Hoerl and Robert W Kennard. 1970 · 1970
Earlier work this paper cites.
Measurement and control of response bias
Delroy L. Paulhus. 1991 · 1991
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Ma teusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Identifying response styles: A latent-class bilinear multinomial logit model
Joost Van Rosmalen, Hester Van Herk, and Patrick J. F. Groenen. 2010 · 2010
Earlier work this paper cites.
Constrained dual scaling for detecting response styles in categorical data
Pieter C. Schoonees, Michel Van de Velden, and Patrick J. F. Groenen. 2015 · 2015
Earlier work this paper cites.
e-SNLI: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Risk-graded safety for handling medical queries in conversational AI
Gavin Abercrombie and Verena Rieser. 2022 · 2022
Earlier work this paper cites.
Zichao Li, Prakhar Sharma, Xing Han Lu, Jackie CK Cheung, and Siva Reddy. 2022 · 2022
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
Perspectives on large language models for relevance judgment
Guglielmo Faggioli, Laura Dietz, Charles L.A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Cited alongside, same era.
Mitigating label biases for in-context learning
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023 · 2023
Cited alongside, same era.
ChatGPT outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Cited alongside, same era.
The Claude 3 model family: Opus, Sonnet, Haiku
Anthropic. 2024 · 2024
Later among the works it cites.
LLMs instead of human judges? A large scale empirical study across 20 NLP evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Later among the works it cites.
Are large language model-based evaluators the solution to scaling up multilingual evaluation?
Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ROSCOE: A suite of metrics for scoring step-by-step reasoning
Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023 · 2023
Cited alongside, same era.
Prototypical calibration for few-shot learning of language models
Zhixiong Han, Yaru Hao, Li Dong, Yutao Sun, and Furu Wei. 2023 · 2023
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Cited alongside, same era.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Automated evaluation of written discourse coherence using GPT-4
Ben Naismith, Phoebe Mulcaire, and Jill Burstein. 2023 · 2023
Cited alongside, same era.
Petter Törnberg. 2023 · 2023
Cited alongside, same era.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023 · 2023
Cited alongside, same era.
ChatGPT rates natural language explanation quality like humans: But on which scales?
Fan Huang, Haewoon Kwak, Kunwoo Park, and Jisun An. 2024 · 2024
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L’elio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Later among the works it cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2024 · 2024
Later among the works it cites.
The effectiveness of LLMs as annotators: A comparative overview and empirical analysis of direct representation
Maja Pavlovic and Massimo Poesio. 2024 · 2024
Later among the works it cites.
Beyond performance: Quantifying and mitigating label bias in LLMs
Yuval Reif and Roy Schwartz. 2024 · 2024
Later among the works it cites.
Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. 2024 · 2024
Later among the works it cites.
Replacing judges with juries: Evaluating LLM generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Later among the works it cites.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2024 · 2024
Later among the works it cites.
Batch calibration: Rethinking calibration for in-context learning and prompt engineering
Han Zhou, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A Heller, and Subhrajit Roy. 2024 · 2024
Later among the works it cites.