Fetching the paper…
Reading the bibliography…
Large language models are often ranked according to their level of alignment with human preferences -- a model is better than other models if its outputs are more frequently preferred by humans.
The USCF Rating System: Its Development, Theory, and Applications
Arpad E. Elo · 1966
Earlier work this paper cites.
League Tables and Their Limitations: Statistical Issues in Comparisons of Institutional Performance
Harvey Goldstein and David J Spiegelhalter · 1996
Earlier work this paper cites.
Reliability of League Tables of In Vitro Fertilisation Clinics: Retrospective Analysis of Live Birth Rates
E Clare Marshall, Colin Sanderson, David J Spiegelhalter, and Martin McKee · 1998
Earlier work this paper cites.
Incorporating Natural Variation into IVF Clinic League Tables
Oscar Lemmers, Jan AM Kremer, and George F Borm · 2007
Earlier work this paper cites.
Using the Bootstrap to Quantify the Authority of an Empirical Ranking
Peter Hall and Hugh Miller · 2009
Earlier work this paper cites.
Confidence Intervals for Population Ranks in the Presence of Ties and Near Ties
Minge Xie, Kesar Singh, and Cun-Hui Zhang · 2009
Earlier work this paper cites.
A Similarity Measure for Indefinite Rankings
William Webber, Alistair Moffat, and Justin Zobel · 2010
Earlier work this paper cites.
Ranking Populations Based on Sample Survey Data
Tommy Wright, Martin Klein, and Jerzy Wieczorek · 2014
Earlier work this paper cites.
Confidence Intervals for Ranks of Age-Adjusted Rates across States or Counties
Shunpu Zhang, Jun Luo, Li Zhu, David G Stinchcomb, Dave Campbell, Ginger Carter, Scott Gilkeson, and Eric J Feuer · 2014
Earlier work this paper cites.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
A Joint Confidence Region for an Overall Ranking of Populations
Martin Klein, Tommy Wright, and Jerzy Wieczorek · 2020
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
A General Language Assistant as a Laboratory for Alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
Justin Rising · 2021
Earlier work this paper cites.
Training Language Models to Follow Instructions with Human Feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Simultaneous Confidence Intervals for Ranks with Application to Ranking Institutions
Diaa Al Mohamad, Jelle J Goeman, and Erik W van Zwet · 2022
Earlier work this paper cites.
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-jian Jiang, and Alexander Rush · 2022
Earlier work this paper cites.
Finetuned Language Models are Zero-Shot Learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le · 2022
Earlier work this paper cites.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2022
Earlier work this paper cites.
Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Earlier work this paper cites.
AI-Generated Medical Advice—GPT and Beyond
Claudia E Haupt and Mason Marks · 2023
Earlier work this paper cites.
Mathematical Discoveries from Program Search with Large Language Models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi · 2023
Cited alongside, same era.
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Aligning Large Language Models with Human: A Survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Elo Uncovered: Robustness and Best Practices in Language Model Evaluation
A Survey on Evaluation of Large Language Models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie · 2024
Closest in time.
Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
LMSYS · 2024
Closest in time.
Stanford Alpaca: An instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2024
Closest in time.
Generative Judge for Evaluating Alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu · 2024
Closest in time.
PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
Ruosen Li, Teerth Patel, and Xinya Du · 2024
Closest in time.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker, and Marzieh Fadaee · 2023
Cited alongside, same era.
Large Language Models Encode Clinical Knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, Blaise Agüera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan · 2023
Cited alongside, same era.
ChatArena: Multi-Agent Language Game Environments for Large Language Models
Yuxiang Wu, Zhengyao Jiang, Akbir Khan, Yao Fu, Laura Ruis, Edward Grefenstette, and Tim Rocktäschel · 2023
Cited alongside, same era.
LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models
Yen-Ting Lin and Yun-Nung Chen · 2023
Cited alongside, same era.
Can Large Language Models Be an Alternative to Human Evaluations?
Cheng-Han Chiang and Hung-yi Lee · 2023
Cited alongside, same era.
LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Cited alongside, same era.
Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang · 2023
Cited alongside, same era.
Preference Proxies: Evaluating Large Language Models in Capturing Human Preferences in Human-AI Tasks
Mudit Verma, Siddhant Bhambri, and Subbarao Kambhampati · 2023
Cited alongside, same era.
Closest in time.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2024
Closest in time.
AutoEval Done Right: Using Synthetic Data for Model Evaluation
Pierre Boyeau, Anastasios N Angelopoulos, Nir Yosef, Jitendra Malik, and Michael I Jordan · 2024
Closest in time.
Large Language Models Can Accurately Predict Searcher Preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra · 2024
Closest in time.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui · 2024
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2024
Closest in time.
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang · 2024
Closest in time.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2024
Closest in time.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky · 2024
Closest in time.
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier · 2024
Closest in time.
Evaluating Language Model Agency Through Negotiations
Tim R Davidson, Veniamin Veselovsky, Martin Josifoski, Maxime Peyrard, Antoine Bosselut, Michal Kosinski, and Robert West · 2024
Closest in time.
Large Language Models are Zero-Shot Rankers for Recommender Systems
Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao · 2024
Closest in time.
Cross-Prediction-Powered Inference
Tijana Zrnic and Emmanuel J Candès · 2024
Closest in time.
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia · 2024
Closest in time.
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
Adian Liusie, Potsawee Manakul, and Mark J. F. Gales · 2024
Closest in time.
Benchmarking Foundation Models with Language-Model-as-an-Examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou · 2024
Closest in time.