Fetching the paper…
Reading the bibliography…
Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models.
Judgment under Uncertainty: Heuristics and Biases
Amos Tversky and Daniel Kahneman · 1974
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
Leo Breiman · 2001
Earlier work this paper cites.
Mining Association Rules for Label Ranking
Cláudio Rebelo de Sá, Carlos Soares, Alípio Mário Jorge, Paulo Azevedo, and Joaquim Costa · 2011
Earlier work this paper cites.
Foundations of Rule Learning
Johannes Fürnkranz, Dragan Gamberger, and Nada Lavrač · 2012
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Training language models to follow instructions with human feedback, March 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Describing Differences between Text Distributions with Natural Language
Ruiqi Zhong, Charlie Snell, Dan Klein, and Jacob Steinhardt · 2022
Earlier work this paper cites.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback, August 2023
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Earlier work this paper cites.
Specific versus General Principles for Constitutional AI, October 2023
Sandipan Kundu, Yuntao Bai, Saurav Kadavath, Amanda Askell, Andrew Callahan, Anna Chen, Anna Goldie, Avital Balwit, Azalia Mirhoseini, Brayden McLean, Catherine Olsson, Cassie Evraets, Eli Tran-Johnson, Esin Durmus, Ethan Perez, Jackson Kernion, Jamie Kerr, Kamal Ndousse, Karina Nguyen, Nelson Elhage, Newton Cheng, Nicholas Schiefer, Nova DasSarma, Oliver Rausch, Robin Larson, Shannon Yang, Shauna Kravec, Timothy Telleen-Lawton, Thomas I. Liao, Tom Henighan, Tristan Hume, Zac Hatfield-Dodds, Sören Mindermann, Nicholas Joseph, Sam McCandlish, and Jared Kaplan · 2023
Earlier work this paper cites.
AlpacaEval: An Automatic Evaluator of Instruction-following Models, May 2024c
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
Calibrating LLM-Based Evaluator, September 2023
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang · 2023
Earlier work this paper cites.
Biases in Large Language Models: Origins, Inventory, and Discussion
Roberto Navigli, Simone Conia, and Björn Ross · 2023
Earlier work this paper cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model, December 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Earlier work this paper cites.
Towards Understanding Sycophancy in Language Models, October 2023
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2023
Cited alongside, same era.
Style Over Substance: Evaluation Biases for Large Language Models, November 2023
Minghao Wu and Alham Fikri Aji · 2023
Cited alongside, same era.
A Critical Evaluation of Evaluations for Long-form Question Answering
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena, July 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale · 2024
Closest in time.
Iterative Interactive Inverse Constitutional AI, 2024
Tim Kostolansky and Julian Manyika · 2024
Closest in time.
Inverse Constitutional AI
Timothy H. Kostolansky · 2024
Closest in time.
Does style matter? Disentangling style and substance in Chatbot Arena, 2024b
Tianle Li, Anastasios Angelopoulos, and Wei-Lin Chiang · 2024
Closest in time.
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs, 2024
Chris Yuhao Liu, Liang Zeng, Jiacai Liu, Rui Yan, Jujie He, Chaojie Wang, Shuicheng Yan, Yang Liu, and Yahui Zhou · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goal Driven Discovery of Distributional Differences via Language Descriptions
Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Dan Klein, and Jacob Steinhardt · 2023
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models, December 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Hritik Bansal, John Dang, and Aditya Grover · 2024
Cited alongside, same era.
Humans or LLMs as the Judge? A Study on Judgement Biases, April 2024
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang · 2024
Cited alongside, same era.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, March 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Cited alongside, same era.
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators, April 2024
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto · 2024
Cited alongside, same era.
Compositional preference models for aligning LMs, March 2024
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, and Marc Dymetman · 2024
Cited alongside, same era.
Human Feedback is not Gold Standard, January 2024
Tom Hosking, Phil Blunsom, and Max Bartolo · 2024
Cited alongside, same era.
Rule Based Rewards for Language Model Safety
Tong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam, Andrea Vallone, Ian D. Kivlichan, Molly Lin, Alex Beutel, John Schulman, and Lilian Weng · 2024
Closest in time.
LLM Evaluators Recognize and Favor Their Own Generations, April 2024
Arjun Panickssery, Samuel R. Bowman, and Shi Feng · 2024
Closest in time.
OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Junsoo Park, Seungyeon Jwa, Ren Meiying, Daeyoung Kim, and Sanghyuk Choi · 2024
Closest in time.
ConstitutionMaker: Interactively Critiquing Large Language Models by Converting Feedback into Principles
Savvas Petridis, Benjamin D Wedin, James Wexler, Mahima Pushkarna, Aaron Donsbach, Nitesh Goyal, Carrie J Cai, and Michael Terry · 2024
Closest in time.
IntentGPT: Few-Shot Intent Discovery with Large Language Models
Juan A. Rodriguez, Nicholas Botzer, David Vazquez, Christopher Pal, Marco Pedersoli, and Issam H. Laradji · 2024
Closest in time.
Large Language Models are Inconsistent and Biased Evaluators, May 2024
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara · 2024
Closest in time.
Cultural bias and cultural alignment of large language models
Yan Tao, Olga Viberg, Ryan S Baker, and René F Kizilcec · 2024
Closest in time.
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment, 2024
Zekun Moore Wang, Shawn Wang, Kang Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, and Wenhao Huang · 2024
Closest in time.