Fetching the paper…
Reading the bibliography…
The rapid advancement of Large Language Models (LLMs) has driven their expanding application across various fields.
The intraclass correlation coefficient as a measure of reliability
John J Bartko. 1966 · 1966
Earlier work this paper cites.
Estimates of the regression coefficient based on Kendall’s tau
Pranab Kumar Sen. 1968 · 1968
Earlier work this paper cites.
Position bias in multiple-choice questions
Niels J Blunch. 1984 · 1984
Earlier work this paper cites.
Evaluations of self and others: Self-enhancement biases in social judgments
Jonathon D Brown. 1986 · 1986
Earlier work this paper cites.
The Turing Test: the first 50 years
Robert M French. 2000 · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Center-of-inattention: Position biases in decision-making
Priya Raghubir and Ana Valenzuela. 2006 · 2006
Earlier work this paper cites.
Pearson correlation coefficient
Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009 · 2009
Earlier work this paper cites.
Computing machinery and intelligence
Alan M Turing. 2009 · 2009
Earlier work this paper cites.
Principles of artificial intelligence
Nils J Nilsson. 2014 · 2014
Earlier work this paper cites.
Spearman’s rank correlation coefficient
Philip Sedgwick. 2014 · 2014
Earlier work this paper cites.
The movielens datasets: History and context
F Maxwell Harper and Joseph A Konstan. 2015 · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Five ways to look at Cohen’s kappa
Matthijs J Warrens. 2015 · 2015
Earlier work this paper cites.
Yelp dataset challenge: Review rating prediction
Nabiha Asghar. 2016 · 2016
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 839–849
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016 · 2016
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of The 8th International Joint Conference on Natural Language Processing (IJCNLP 2017)
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Position bias estimation for unbiased learning to rank in personal search. In Proceedings of the eleventh ACM international conference on web search and data mining . 610–618
Xuanhui Wang, Nadav Golbandi, Michael Bendersky, Donald Metzler, and Marc Najork. 2018 · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too
Saizheng Zhang. 2018 · 2018
Earlier work this paper cites.
Look at the first sentence: Position bias in question answering
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Earlier work this paper cites.
Natural backdoor attack on text data
Lichao Sun. 2020 · 2020
Earlier work this paper cites.
Adv-bert: Bert is not robust on misspellings! generating nature adversarial samples on bert
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, and Caiming Xiong. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment. arXiv
A Askell, Y Bai, A Chen, D Drain, D Ganguli, T Henighan, A Jones, N Joseph, B Mann, N DasSarma, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Overview of the TREC 2021 Deep Learning Track. In TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2021 · 2021
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a · 2021
Earlier work this paper cites.
OpenMEVA: A benchmark for evaluating open-ended story generation metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021 · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Earlier work this paper cites.
Automatic evaluation and moderation of open-domain dialogue systems
Chen Zhang, João Sedoc, Luis Fernando D’Haro, Rafael Banchs, and Alexander Rudnicky. 2021 · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models. In International conference on machine learning . PMLR, 12697–12706
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples
Hezekiah J Branch, Jonathan Rodriguez Cefalu, Jeremy McHugh, Leyla Hujer, Aditya Bahl, Daniel del Castillo Iglesias, Ron Heichman, and Ramesh Darwishi. 2022 · 2022
Earlier work this paper cites.
Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Cyril Chhun, Pierre Colombo, Chloé Clavel, and Fabian M Suchanek. 2022 · 2022
Earlier work this paper cites.
Overview of the TREC 2022 Deep Learning Track. In TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2022 · 2022
Earlier work this paper cites.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, et al · 2022
Earlier work this paper cites.
Prototypical calibration for few-shot learning of language models
Zhixiong Han, Yaru Hao, Li Dong, Yutao Sun, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
Query-efficient and scalable black-box adversarial attacks on discrete sequential data via bayesian optimization. In International Conference on Machine Learning . PMLR, 12478–12497
Deokjae Lee, Seungyong Moon, Junhyeok Lee, and Hyun Oh Song. 2022 · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023 · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Earlier work this paper cites.
StoryER: Automatic story evaluation via ranking, rating and reasoning
Hong Chen, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao, and Hideki Nakayama. 2023b · 2023
Earlier work this paper cites.
Adaptation with self-evaluation to improve selective prediction in llms
Jiefeng Chen, Jinsung Yoon, Sayna Ebrahimi, Sercan O Arik, Tomas Pfister, and Somesh Jha. 2023c · 2023
Earlier work this paper cites.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023a · 2023
Earlier work this paper cites.
A closer look into automatic evaluation using large language models
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Earlier work this paper cites.
Mitigating label biases for in-context learning
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023 · 2023
Earlier work this paper cites.
Knowledge solver: Teaching llms to search for domain knowledge from knowledge graphs
Chao Feng, Xinyu Zhang, and Zichu Fei. 2023 · 2023
Earlier work this paper cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2023 · 2023
Earlier work this paper cites.
ChatGPT outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Earlier work this paper cites.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tur. 2023 · 2023
Earlier work this paper cites.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al · 2023
Earlier work this paper cites.
Are large language model-based evaluators the solution to scaling up multilingual evaluation?
Rishav Hada, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2023 · 2023
Earlier work this paper cites.
ALLURE: auditing and improving llm-based evaluation of text using iterative in-context-learning
Hosein Hasanbeig, Hiteshi Sharma, Leo Betthauser, Felipe Vieira Frujeri, and Ida Momennejad. 2023 · 2023
Earlier work this paper cites.
Socreval: Large language models with the socratic method for reference-free reasoning evaluation
Hangfeng He, Hongming Zhang, and Dan Roth. 2023b · 2023
Earlier work this paper cites.
Annollm: Making large language models to be better crowdsourced annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, Alex Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al · 2023
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023 · 2023
Earlier work this paper cites.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023 · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
Tigerscore: Towards building explainable metric for all text generation tasks
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023b · 2023
Earlier work this paper cites.
Prompt packer: Deceiving llms through compositional instruction with hidden attacks
Shuyu Jiang, Xingshu Chen, and Rui Tang. 2023a · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?, 2024
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Marzena Karpinska and Mohit Iyyer. 2023 · 2023
Earlier work this paper cites.
Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al · 2023
Earlier work this paper cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Earlier work this paper cites.
Neema Kotonya, Saran Krishnasamy, Joel Tetreault, and Alejandro Jaimes. 2023 · 2023
Earlier work this paper cites.
Can large language models aid in annotating speech emotional data? uncovering new frontiers
Siddique Latif, Muhammad Usama, Mohammad Ibrahim Malik, and Björn W Schuller. 2023 · 2023
Earlier work this paper cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi. 2023 · 2023
Earlier work this paper cites.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023c · 2023
Earlier work this paper cites.
Qintong Li, Leyang Cui, Lingpeng Kong, and Wei Bi. 2023a · 2023
Earlier work this paper cites.
Prd: Peer rank and discussion improve large language model based evaluations
Ruosen Li, Teerth Patel, and Xinya Du. 2023b · 2023
Earlier work this paper cites.
Split and merge: Aligning position biases in large language model based evaluators
Zongjie Li, Chaozheng Wang, Pingchuan Ma, Daoyuan Wu, Shuai Wang, Cuiyun Gao, and Yang Liu. 2023d · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023 · 2023
Earlier work this paper cites.
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Earlier work this paper cites.
Minqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, and Lifu Huang. 2023c · 2023
Earlier work this paper cites.
Alignbench: Benchmarking chinese alignment of large language models
Xiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang, Zhuoer Feng, Bosi Wen, Jiale Cheng, Pei Ke, Yifan Xu, Weng Lam Tam, et al · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Earlier work this paper cites.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a · 2023
Earlier work this paper cites.
Calibrating llm-based evaluator
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2023d · 2023
Earlier work this paper cites.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. 2023 · 2023
Earlier work this paper cites.
BioPlanner: automatic evaluation of LLMs on protocol planning in biology
Odhran O’Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud, Ali Essa Ghareeb, Justin Booth, and Samuel G Rodriques. 2023 · 2023
Earlier work this paper cites.
Refiner: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings. 2023 · 2023
Earlier work this paper cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. 2023 · 2023
Earlier work this paper cites.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al · 2023
Earlier work this paper cites.
Self-evaluation improves selective generation in large language models. In Proceedings on . PMLR, 49–64
Jie Ren, Yao Zhao, Tu Vu, Peter J Liu, and Balaji Lakshminarayanan. 2023 · 2023
Earlier work this paper cites.
Retrieval-based Evaluation for LLMs: A Case Study in Korean Legal QA. In Proceedings of the Natural Legal Language Processing Workshop 2023 . 132–137
Cheol Ryu, Seolhwa Lee, Subeen Pang, Chanyeol Choi, Hojun Choi, Myeonggee Min, and Jy-Yong Sohn. 2023 · 2023
Earlier work this paper cites.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Cited alongside, same era.
Opinsummeval: Revisiting automated evaluation for opinion summarization
Yuchen Shen and Xiaojun Wan. 2023 · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning . PMLR, 31210–31227
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
RecExplainer: Aligning Large Language Models for Explaining Recommendation Models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1530–1541
Yuxuan Lei, Jianxun Lian, Jing Yao, Xu Huang, Defu Lian, and Xing Xie. 2024 · 2024
Closest in time.
Automatic evaluation for mental health counseling using llms
Anqi Li, Yu Lu, Nirui Song, Shuai Zhang, Lizhi Ma, and Zhenzhong Lan. 2024c · 2024
Closest in time.
Calibraeval: Calibrating prediction distribution to mitigate selection bias in llms-as-judges
Haitao Li, Junjie Chen, Qingyao Ai, Zhumin Chu, Yujia Zhou, Qian Dong, and Yiqun Liu. 2024a · 2024
Closest in time.
Debatrix: Multi-dimensinal Debate Judge with Iterative Chronological Analysis Based on LLM
Jingcong Liang, Rong Ye, Meng Han, Ruofei Lai, Xinyu Zhang, Xuanjing Huang, and Zhongyu Wei. 2024a · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Avi Singh, John D Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J Liu, James Harrison, Jaehoon Lee, Kelvin Xu, et al · 2023
Cited alongside, same era.
Andrea Sottana, Bin Liang, Kai Zou, and Zheng Yuan. 2023 · 2023
Cited alongside, same era.
Petter Törnberg. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
LLMs cannot find reasoning errors, but can correct them!
Gladys Tyen, Hassan Mansoor, Peter Chen, Tony Mak, and Victor Cărbune. 2023 · 2023
Cited alongside, same era.
Can large language models really improve by self-critiquing their own plans?
Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati. 2023 · 2023
Cited alongside, same era.
Learning Evaluation Models from Large Language Models for Sequence Generation
Chenglong Wang, Hang Zhou, Kaiyan Chang, Tongran Liu, Chunliang Zhang, Quan Du, Tong Xiao, and Jingbo Zhu. 2023e · 2023
Cited alongside, same era.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023b · 2023
Cited alongside, same era.
ABSEval: An Agent-based Framework for Script Evaluation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 12418–12434
Sirui Liang, Baoli Zhang, Jun Zhao, and Kang Liu. 2024b · 2024
Closest in time.
WILDBENCH: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. 2024 · 2024
Closest in time.
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024a · 2024
Closest in time.
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, and Juanzi Li. 2024b · 2024
Closest in time.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier. 2024c · 2024
Closest in time.
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
Adian Liusie, Vatsal Raina, Yassir Fathullah, and Mark Gales. 2024 · 2024
Closest in time.
Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, and Junnan Li. 2024 · 2024
Closest in time.
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
Shengjie Ma, Chong Chen, Qi Chu, and Jiaxin Mao. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Closest in time.
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
John Mendonça, Isabel Trancoso, and Alon Lavie. 2024 · 2024
Closest in time.
Evaluating the Performance of Large Language Models via Debates
Behrad Moniri, Hamed Hassani, and Edgar Dobriban. 2024 · 2024
Closest in time.
Creative Beam Search: LLM-as-a-Judge for Improving Response Generation. ICCC
Mirco Musolesi. 2024 · 2024
Closest in time.
Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen. 2024 · 2024
Closest in time.
JurEE not Judges: safeguarding llm interactions with small, specialised Encoder Ensembles
Dom Nasrabadi. 2024 · 2024
Closest in time.
PiCO: Peer Review in LLMs based on the Consistency Optimization
Kun-Peng Ning, Shuo Yang, Yuyang Liu, Jia-Yu Yao, Zhenhui Liu, Yu Wang, Ming Pang, and Li Yuan. 2024 · 2024
Closest in time.
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
Tong Niu, Shafiq Joty, Ye Liu, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. 2024 · 2024
Closest in time.
A multi-llm debiasing framework
Deonna M Owens, Ryan A Rossi, Sungchul Kim, Tong Yu, Franck Dernoncourt, Xiang Chen, Ruiyi Zhang, Jiuxiang Gu, Hanieh Deilamsalehy, and Nedim Lipka. 2024 · 2024
Closest in time.
Human-Centered Design Recommendations for LLM-as-a-judge
Qian Pan, Zahra Ashktorab, Michael Desmond, Martin Santillan Cooper, James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer. 2024a · 2024
Closest in time.
Unifying large language models and knowledge graphs: A roadmap
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024b · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng. 2024 · 2024
Closest in time.
Offsetbias: Leveraging debiased data for tuning evaluators
Junsoo Park, Seungyeon Jwa, Meiying Ren, Daeyoung Kim, and Sanghyuk Choi. 2024 · 2024
Closest in time.
AIME: AI System Optimization via Multiple LLM Evaluators
Bhrij Patel, Souradip Chakraborty, Wesley A Suttle, Mengdi Wang, Amrit Singh Bedi, and Dinesh Manocha. 2024 · 2024
Closest in time.
Bias patterns in the application of LLMs for clinical decision support: A comprehensive study
Raphael Poulain, Hamed Fayyaz, and Rahmatollah Beheshti. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
LLMJudge: LLMs for Relevance Judgments
Hossein A Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas, Charles LA Clarke, Mohammad Aliannejadi, Clemencia Siro, and Guglielmo Faggioli. 2024 · 2024
Closest in time.
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Vyas Raina, Adian Liusie, and Mark Gales. 2024 · 2024
Closest in time.
Constructing domain-specific evaluation sets for llm-as-a-judge
Ravi Raju, Swayambhoo Jain, Bo Li, Jonathan Li, and Urmish Thakker. 2024 · 2024
Closest in time.
A systematic survey of prompt engineering in large language models: Techniques and applications
Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha. 2024 · 2024
Closest in time.
Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–14
Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024 · 2024
Closest in time.
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024b · 2024
Closest in time.
Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi. 2024a · 2024
Closest in time.
Fusion-Eval: Integrating Assistant Evaluators with LLMs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track . 225–238
Lei Shu, Nevan Wichers, Liangchen Luo, Yun Zhu, Yinxiao Liu, Jindong Chen, and Lei Meng. 2024 · 2024
Closest in time.
Don’t Use LLMs to Make Relevance Judgments
Ian Soboroff. 2024 · 2024
Closest in time.
KRX Bench: Automating Financial Benchmark Creation via Large Language Models. In Proceedings of the Joint Workshop of the 7th Financial Technology and Natural Language Processing, the 5th Knowledge Discovery from Unstructured Data in Financial Services, and the 4th Workshop on Economics and Natural Language Processing@ LREC-COLING 2024 . 10–20
Guijin Son, Hyunjun Jeon, Chami Hwang, and Hanearl Jung. 2024a · 2024
Closest in time.
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
Guijin Son, Dongkeun Yoon, Juyoung Suk, Javier Aula-Blasco, Mano Aslan, Vu Trong Kim, Shayekh Bin Islam, Jaume Prats-Cristià, Lucía Tormo-Bañuelos, and Seungone Kim. 2024b · 2024
Closest in time.
FineSurE: Fine-grained summarization evaluation using LLMs
Hwanjun Song, Hang Su, Igor Shalyminov, Jason Cai, and Saab Mansour. 2024a · 2024
Closest in time.
Can Many-Shot In-Context Learning Help Long-Context LLM Judges? See More, Judge Better!
Mingyang Song, Mao Zheng, and Xuan Luo. 2024b · 2024
Closest in time.
Automated Essay Scoring and Revising Based on Open-Source Large Language Models
Yishen Song, Qianta Zhu, Huaibo Wang, and Qinhua Zheng. 2024c · 2024
Closest in time.
From calculation to adjudication: Examining llm judges on mathematical reasoning tasks
Andreas Stephan, Dawei Zhu, Matthias Aßenmacher, Xiaoyu Shen, and Benjamin Roth. 2024 · 2024
Closest in time.
Large language models are inconsistent and biased evaluators
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Closest in time.
Fast Best-of-N Decoding via Speculative Rejection
Hanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang, Jiahao Qiu, Ming Yin, Mengdi Wang, Peter Bartlett, and Andrea Zanette. 2024 · 2024
Closest in time.
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Annalisa Szymanski, Noah Ziems, Heather A Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer. 2024 · 2024
Closest in time.
JudgeBench: A Benchmark for Evaluating LLM-based Judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024 · 2024
Closest in time.
AI can help humans find common ground in democratic deliberation
Michael Henry Tessler, Michiel A Bakker, Daniel Jarrett, Hannah Sheahan, Martin J Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C Parkes, et al · 2024
Closest in time.
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024 · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024 · 2024
Closest in time.
Appworld: A controllable world of apps and people for benchmarking interactive coding agents
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian. 2024b · 2024
Closest in time.
Self-rationalization improves LLM as a fine-grained judge
Prapti Trivedi, Aditya Gulati, Oliver Molenschot, Meghana Arakkal Rajeev, Rajkumar Ramamurthy, Keith Stevens, Tanveesh Singh Chaudhery, Jahnavi Jambholkar, James Zou, and Nazneen Rajani. 2024a · 2024
Closest in time.
Are Expert-Level Language Models Expert-Level Annotators?
Yu-Min Tseng, Wei-Lin Chen, Chung-Chi Chen, and Hsin-Hsi Chen. 2024 · 2024
Closest in time.
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Closest in time.
Foundational autoraters: Taming large language models for better automatic evaluation
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, and Yun-Hsuan Sung. 2024 · 2024
Closest in time.
Halu-j: Critique-based hallucination judge
Binjie Wang, Steffi Chern, Ethan Chern, and Pengfei Liu. 2024b · 2024
Closest in time.
Automated Genre-Aware Article Scoring and Feedback Using Large Language Models
Chihang Wang, Yuxin Dong, Zhenhong Zhang, Ruotong Wang, Shuo Wang, and Jiajing Chen. 2024c · 2024
Closest in time.
Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li. 2024e · 2024
Closest in time.
Wanying Wang, Zeyu Ma, Pengfei Liu, and Mingang Chen. 2024f · 2024
Closest in time.
HelpSteer2-Preference: Complementing Ratings with Preferences
Zhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert, Gerald Shen, Jiaqi Zeng, Oleksii Kuchaiev, and Yi Dong. 2024a · 2024
Closest in time.
Cream: Consistency regularized self-rewarding language models
Zhaoyang Wang, Weilei He, Zhiyuan Liang, Xuchao Zhang, Chetan Bansal, Ying Wei, Weitong Zhang, and Huaxiu Yao. 2024d · 2024
Closest in time.
Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram. 2024 · 2024
Closest in time.
Martin Weyssow, Aton Kamanda, and Houari Sahraoui. 2024 · 2024
Closest in time.
Continual learning for large language models: A survey
Tongtong Wu, Linhao Luo, Yuan-Fang Li, Shirui Pan, Thuy-Trang Vu, and Gholamreza Haffari. 2024a · 2024
Closest in time.
Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge
Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar. 2024b · 2024
Closest in time.
Evaluating Mathematical Reasoning Beyond Accuracy
Shijie Xia, Xuefeng Li, Yixin Liu, Tongshuang Wu, and Pengfei Liu. 2024a · 2024
Closest in time.
Language Models can Evaluate Themselves via Probability Discrepancy
Tingyu Xia, Bowen Yu, Yuan Wu, Yi Chang, and Chang Zhou. 2024b · 2024
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al · 2024
Closest in time.
Self-evaluation guided beam search for reasoning
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao, Min-Yen Kan, Junxian He, and Michael Xie. 2024a · 2024
Closest in time.
Improving Model Factuality with Fine-grained Critique-based Evaluator
Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Han Fang, Carolyn Rose, et al · 2024
Closest in time.
Llava-critic: Learning to evaluate multimodal models
Tianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye, Haoqi Fan, Quanquan Gu, Heng Huang, and Chunyuan Li. 2024 · 2024
Closest in time.
Large Language Models Are Active Critics in NLG Evaluation
Shuying Xu, Junjie Hu, and Ming Jiang. 2024b · 2024
Closest in time.
The perfect blend: Redefining RLHF with mixture of judges
Tengyu Xu, Eryk Helenowski, Karthik Abinav Sankararaman, Di Jin, Kaiyan Peng, Eric Han, Shaoliang Nie, Chen Zhu, Hejia Zhang, Wenxuan Zhou, et al · 2024
Closest in time.
Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing
Le Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman, Xuanhui Wang, Michael Bendersky, and Harrie Oosterhuis. 2024a · 2024
Closest in time.
Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 410–423
Le Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman, Xuanhui Wang, Michael Bendersky, and Harrie Oosterhuis. 2024b · 2024
Closest in time.
Mitigating biases for instruction-following language models via bias neurons elimination. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 9061–9073
Nakyeong Yang, Taegwan Kang, Stanley Jungkyu Choi, Honglak Lee, and Kyomin Jung. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Self-Judge: Selective Instruction Following with Alignment Self-Evaluation
Hai Ye and Hwee Tou Ng. 2024 · 2024
Closest in time.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al · 2024
Closest in time.
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou, Wei Shen, Dong Yan, and Yiqun Liu. 2024a · 2024
Closest in time.
Seungjun Yi, Jaeyoung Lim, and Juyong Yoon. 2024 · 2024
Closest in time.
Kieval: A knowledge-grounded interactive evaluation framework for large language models
Zhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang, Wei Ye, Jindong Wang, Xing Xie, Yue Zhang, and Shikun Zhang. 2024 · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.
STaR: Self-taught reasoner bootstrapping reasoning with reasoning. In Proc. the 36th International Conference on Neural Information Processing Systems , Vol. 1126
Eric Zelikman, YH Wu, Jesse Mu, and Noah D Goodman. 2024 · 2024
Closest in time.
Automatic Instruction Evolving for Large Language Models
Weihao Zeng, Can Xu, Yingxiu Zhao, Jian-Guang Lou, and Weizhu Chen. 2024 · 2024
Closest in time.
Kaiqi Zhang, Shuai Yuan, and Honghan Zhao. 2024c · 2024
Closest in time.
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
Qiyuan Zhang, Yufei Wang, Tiezheng Yu, Yuxin Jiang, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, et al · 2024
Closest in time.
Auto Arena of LLMs: Automating LLM Evaluations with Agent Peer-battles and Committee Discussions
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Deli Zhao, and Lidong Bing. 2024b · 2024
Closest in time.
Measuring the inconsistency of large language models in preferential ranking
Xiutian Zhao, Ke Wang, and Wei Peng. 2024a · 2024
Closest in time.
Cheating automatic llm benchmarks: Null models achieve high win rates
Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin. 2024 · 2024
Closest in time.
Mitigating the Bias of Large Language Model Evaluation
Hongli Zhou, Hui Huang, Yunfei Long, Bing Xu, Conghui Zhu, Hailong Cao, Muyun Yang, and Tiejun Zhao. 2024c · 2024
Closest in time.
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
Han Zhou, Xingchen Wan, Yinhong Liu, Nigel Collier, Ivan Vulić, and Anna Korhonen. 2024e · 2024
Closest in time.
Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing Tasks. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . 9340–9351
Ruiyang Zhou, Lu Chen, and Kai Yu. 2024a · 2024
Closest in time.
Calibrated self-rewarding vision language models
Yiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang, Zhaorun Chen, Chenhang Cui, Xiyao Wang, Yun Li, Linjun Zhang, and Huaxiu Yao. 2024b · 2024
Closest in time.
Are Large Language Models Rational Investors?
Yuhang Zhou, Yuchen Ni, Xiang Liu, Jian Zhang, Sen Liu, Guangnan Ye, and Hongfeng Chai. 2024d · 2024
Closest in time.
A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24) . Association for Computing Machinery, New York, NY, USA, 38–47
Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. 2024 · 2024
Closest in time.
Agent-as-a-Judge: Evaluate Agents with Agents
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al · 2024
Closest in time.
JUDING THE JUDGES: ASYSTEMATIC INVESTIGATION OF POSITION BIAS IN PAIRWISE COMPARATIVE AS
SESSMENTS BY LLMS. 2025 · 2025
Closest in time.