Fetching the paper…
Reading the bibliography…
The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 1903
Earlier work this paper cites.
The effect of forced choice on choice
Ravi Dhar and Itamar Simonson · 2003
Earlier work this paper cites.
The physics of optimal decision making: a formal analysis of models of performance in two-alternative forced-choice tasks
Rafal Bogacz, Eric Brown, Jeff Moehlis, Philip Holmes, and Jonathan D Cohen · 2006
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Earlier work this paper cites.
A bayesian framework for modeling human evaluations
Himabindu Lakkaraju, Jure Leskovec, Jon Kleinberg, and Sendhil Mullainathan · 2015
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
Modelling forced-choice response formats
Anna Brown and Albert Maydeu-Olivares · 2018
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2019
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski · 2019
Earlier work this paper cites.
Human uncertainty makes classification more robust
Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Russakovsky · 2019
Earlier work this paper cites.
Learning from noisy labels by regularized estimation of annotator confusion
Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C Alexander, and Nathan Silberman · 2019
Earlier work this paper cites.
Efficient conformal prediction via cascaded inference with expanded admission
Adam Fisch, Tal Schuster, Tommi Jaakkola, and Regina Barzilay · 2020
Earlier work this paper cites.
Ambigqa: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2020
Earlier work this paper cites.
What can we learn from collective human opinions on natural language inference data?
Yixin Nie, Xiang Zhou, and Mohit Bansal · 2020
Earlier work this paper cites.
Would you describe a leopard as yellow? evaluating crowd-annotations with justified and informative disagreement
Pia Sommerauer, Antske Fokkens, and Piek Vossen · 2020
Earlier work this paper cites.
A case for soft loss functions
Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis · 2020
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Earlier work this paper cites.
Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning
Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio, et al · 2021
Earlier work this paper cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Mitchell L Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S Bernstein · 2021
Earlier work this paper cites.
Eliciting and learning with soft labels from every annotator
Katherine M Collins, Umang Bhatt, and Adrian Weller · 2022
Earlier work this paper cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Jury learning: Integrating dissenting voices into machine learning models
Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein · 2022
Earlier work this paper cites.
Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation
Nitesh Goyal, Ian D Kivlichan, Rachel Rosen, and Lucy Vasserman · 2022
Cited alongside, same era.
The’problem’of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank · 2022
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
John Schulman, Barret Zoph, Christina Kim, Jacob Hilton, Jacob Menick, Jiayi Weng, Juan Felipe Ceron Uribe, Liam Fedus, Luke Metz, Michael Pokorny, et al · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Toward a perspectivist turn in ground truthing for predictive computing
Federico Cabitza, Andrea Campagner, and Valerio Basile · 2023
Perspectivist approaches to natural language processing: a survey
Simona Frenda, Gavin Abercrombie, Valerio Basile, Alessandro Pedrani, Raffaella Panizzon, Alessandra Teresa Cignarella, Cristina Marco, and Davide Bernardi · 2024
Later among the works it cites.
Bayesian calibration of win rate estimation with llm evaluators
Yicheng Gao, Gonghan Xu, Zhe Wang, and Arman Cohan · 2024
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al · 2024
Later among the works it cites.
Trust or escalate: Llm judges with provable guarantees for human agreement
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement
Quan Ze Chen and Amy X Zhang · 2023
Cited alongside, same era.
Ragas: Automated evaluation of retrieval augmented generation
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert · 2023
Cited alongside, same era.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tur · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Annotation error detection: Analyzing the past and present for a more coherent future
Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych · 2023
Cited alongside, same era.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka · 2023
Cited alongside, same era.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2023
Cited alongside, same era.
Jaehun Jung, Faeze Brahman, and Yejin Choi · 2024
Later among the works it cites.
Evallm: Interactive evaluation of large language model prompts on user-defined criteria
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim · 2024
Later among the works it cites.
Llm-mod: Can large language models assist content moderation?
Mahi Kolla, Siddharth Salunkhe, Eshwar Chandrasekharan, and Koustuv Saha · 2024
Later among the works it cites.
Split and merge: Aligning position biases in llm-based evaluators
Zongjie Li, Chaozheng Wang, Pingchuan Ma, Daoyuan Wu, Shuai Wang, Cuiyun Gao, and Yang Liu · 2024
Later among the works it cites.
Can vision-language models replace human annotators: A case study with celeba dataset
Haoming Lu and Feifei Zhong · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al · 2024
Later among the works it cites.
Are large language models reliable argument quality annotators?
Nailia Mirzakhmedova, Marcel Gohsen, Chia Hao Chang, and Benno Stein · 2024
Later among the works it cites.
Llmjudge: Llms for relevance judgments
Hossein A Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas, Charles LA Clarke, Mohammad Aliannejadi, Clemencia Siro, and Guglielmo Faggioli · 2024
Later among the works it cites.
Finding replicable human evaluations via stable ranking probability
Parker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, and Markus Freitag · 2024
Later among the works it cites.
Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences
Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya G Parameswaran, and Ian Arawjo · 2024
Later among the works it cites.
Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi · 2024
Later among the works it cites.
Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks
Annalisa Szymanski, Noah Ziems, Heather A Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer · 2024
Later among the works it cites.
Llm-assisted relevance assessments: When should we ask llms for help?
Rikiya Takehi, Ellen M Voorhees, and Tetsuya Sakai · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica · 2024
Later among the works it cites.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes · 2024
Later among the works it cites.
Know your limits: A survey of abstention in large language models
Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al · 2024
Later among the works it cites.
Nishant Balepur, Rachel Rudinger, and Jordan Lee Boyd-Graber · 2025
Closest in time.
Sources of disagreement in data for llm instruction tuning
Russel Dsouza and Venelin Kovatchev · 2025
Closest in time.
Context and Meaning: Navigating Disagreements in NLP Annotation , 2025. Association for Computational Linguistics
Michael Roth and Dominik Schlechtweg, editors · 2025
Closest in time.