Fetching the paper…
Reading the bibliography…
Recent work on the limitations of using reinforcement learning from human feedback (RLHF) to incorporate human preferences into model behavior often raises social choice theory as a reference point.
Deliberative democracy and social choice
David Miller · 1992
Earlier work this paper cites.
The distortion of cardinal preferences in voting
Ariel D Procaccia and Jeffrey S Rosenschein · 2006
Earlier work this paper cites.
Sample complexity for winner prediction in elections
Palash Dey and Arnab Bhattacharyya · 2015
Earlier work this paper cites.
Handbook of computational social choice
Felix Brandt, Vincent Conitzer, Ulle Endriss, Jérôme Lang, and Ariel D Procaccia · 2016
Earlier work this paper cites.
Learning mixtures of plackett-luce models
Zhibing Zhao, Peter Piech, and Lirong Xia · 2016
Earlier work this paper cites.
Social choice under metric preferences: Scoring rules and stv
Piotr Skowron and Edith Elkind · 2017
Earlier work this paper cites.
Approximating optimal social choice under metric preferences
Elliot Anshelevich, Onkar Bhardwaj, Edith Elkind, John Postl, and Piotr Skowron · 2018
Earlier work this paper cites.
Efficiently learning mixtures of mallows models
Allen Liu and Ankur Moitra · 2018
Earlier work this paper cites.
Learning mixtures of random utility models
Zhibing Zhao, Tristan Villamil, and Lirong Xia · 2018
Earlier work this paper cites.
Low-distortion social welfare functions
Gerdus Benade, Ariel D. Procaccia, and Mingda Qiao · 2019
Earlier work this paper cites.
Stretching the effectiveness of mle from accuracy to bias for pairwise comparisons
Jingyan Wang, Nihar Shah, and R Ravi · 2020
Cited alongside, same era.
Distortion in social choice problems: The first 15 years and beyond
Elliot Anshelevich, Aris Filos-Ratsikas, Nisarg Shah, and Alexandros A Voudouris · 2021
Cited alongside, same era.
Learning reward functions from scale feedback
Nils Wilde, Erdem Biyik, Dorsa Sadigh, and Stephen L Smith · 2022
Cited alongside, same era.
Dices dataset: Diversity in conversational ai evaluation for safety, 2023
Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
The history and risks of reinforcement learning and human feedback
Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick · 2023
Later among the works it cites.
Statistical rejection sampling improves preference optimization
Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu · 2023
Later among the works it cites.
Ai alignment and social choice: Fundamental limitations and policy implications
Abhilash Mishra · 2023
Later among the works it cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sara Fish, Paul Gölz, David C Parkes, Ariel D Procaccia, Gili Rusak, Itai Shapira, and Manuel Wüthrich · 2023
Cited alongside, same era.
Smoothed analysis of social choice revisited
Bailey Flanigan, Daniel Halpern, and Alexandros Psomas · 2023
Cited alongside, same era.
Representation with incomplete votes
Daniel Halpern, Gregory Kehne, Ariel D Procaccia, Jamie Tucker-Foltz, and Manuel Wüthrich · 2023
Cited alongside, same era.
Algorithmic collective action in machine learning
Moritz Hardt, Eric Mazumdar, Celestine Mendler-Dünner, and Tijana Zrnic · 2023
Cited alongside, same era.
Incorporating worker perspectives into mturk annotation practices for nlp
Olivia Huang, Eve Fleisig, and Dan Klein · 2023
Cited alongside, same era.
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
A long way to go: Investigating length correlations in rlhf
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2023
Later among the works it cites.
Distributional preference learning: Understanding and accounting for hidden context in rlhf
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons, 2023
Banghua Zhu, Jiantao Jiao, and Michael I. Jordan · 2023
Later among the works it cites.
Social choice for ai alignment: Dealing with diverse human feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, and William S. Zwicker · 2024
Closest in time.