Fetching the paper…
Reading the bibliography…
After pre-training, large language models are aligned with human preferences based on pairwise comparisons.
Probabilistic social choice based on simple voting comparisons
Peter C. Fishburn · 1984
Earlier work this paper cites.
Common voting rules as maximum likelihood estimators
Vincent Conitzer and Tuomas Sandholm · 2005
Earlier work this paper cites.
The distortion of cardinal preferences in voting
Ariel D. Procaccia and Jeffrey S. Rosenschein · 2006
Earlier work this paper cites.
Voting almost maximizes social welfare despite limited communication
Ioannis Caragiannis and Ariel D. Procaccia · 2011
Earlier work this paper cites.
Optimal social choice functions: A utilitarian view
Craig Boutilier, Ioannis Caragiannis, Simi Haber, Tyler Lu, Ariel D. Procaccia, and Or Sheffet · 2012
Earlier work this paper cites.
Statistical methods for ranking data , volume 1341
Mayer Alvo and LH Philip · 2014
Earlier work this paper cites.
A statistical decision-theoretic framework for social choice
Hossein Azari Soufiani, David C Parkes, and Lirong Xia · 2014
Earlier work this paper cites.
Learning mixtures of plackett-luce models
Zhibing Zhao, Peter Piech, and Lirong Xia · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Approximating optimal social choice under metric preferences
Elliot Anshelevich, Onkar Bhardwaj, Edith Elkind, John Postl, and Piotr Skowron · 2018
Earlier work this paper cites.
Bayesian estimators as voting rules
Lirong Xia · 2018
Earlier work this paper cites.
Learning and decision-making from rank data
Lirong Xia · 2019
Earlier work this paper cites.
Learning mixtures of plackett-luce models from structured partial orders
Zhibing Zhao and Lirong Xia · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Resolving the optimal metric distortion conjecture
Vasilis Gkatzelis, Daniel Halpern, and Nisarg Shah · 2020
Earlier work this paper cites.
Axioms for learning from pairwise comparisons
Ritesh Noothigattu, Dominik Peters, and Ariel D Procaccia · 2020
Earlier work this paper cites.
Distortion in social choice problems: The first 15 years and beyond
Elliot Anshelevich, Aris Filos-Ratsikas, Nisarg Shah, and Alexandros A. Voudouris · 2021
Earlier work this paper cites.
Preference elicitation for participatory budgeting
Gerdus Benade, Swaprava Nath, Ariel D. Procaccia, and Nisarg Shah · 2021
Cited alongside, same era.
Plurality Veto: A Simple Voting Rule Achieving Optimal Metric Distortion
Fatih Erdem Kizilkaya and David Kempe · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
On the identifiability of mixtures of ranking models
Xiaomin Zhang, Xucheng Zhang, Po-Ling Loh, and Yingyu Liang · 2022
Cited alongside, same era.
Distortion Under Public-Spirited Voting
Bailey Flanigan, Ariel D Procaccia, and Sven Wang · 2023
Cited alongside, same era.
Ai alignment and social choice: Fundamental limitations and policy implications
KTO: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Later among the works it cites.
Axioms for AI alignment from human feedback
Luise Ge, Daniel Halpern, Evi Micha, Ariel D Procaccia, Itai Shapira, Yevgeniy Vorobeychik, and Junlin Wu · 2024
Later among the works it cites.
Audrey Huang, Wenhao Zhan, Tengyang Xie, Jason D Lee, Wen Sun, Akshay Krishnamurthy, and Dylan J Foster · 2024
Later among the works it cites.
Provably mitigating overoptimization in RLHF: Your SFT loss is implicitly an adversarial regularizer
Zhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu, Hongyi Guo, Yingxiang Yang, Jose Blanchet, and Zhaoran Wang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abhilash Mishra · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Distributional preference learning: Understanding and accounting for hidden context in RLHF
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2023
Cited alongside, same era.
Is RLHF more difficult than standard RL? a theoretical perspective
Yuanhao Wang, Qinghua Liu, and Chi Jin · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Cited alongside, same era.
Human alignment of large language models through online preference optimisation
Daniele Calandriello, Zhaohan Daniel Guo, Remi Munos, Mark Rowland, Yunhao Tang, Bernardo Avila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, et al · 2024
Cited alongside, same era.
Maxmin-RLHF: Alignment with diverse human preferences
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Singh Bedi, and Mengdi Wang · 2024
Cited alongside, same era.
Yu Meng, Mengzhou Xia, and Danqi Chen · 2024
Later among the works it cites.
Nash learning from human feedback
Remi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Côme Fiegel, et al · 2024
Later among the works it cites.
RLHF from heterogeneous feedback via personalization and preference aggregation
Chanwoo Park, Mingyang Liu, Dingwen Kong, Kaiqing Zhang, and Asuman Ozdaglar · 2024
Later among the works it cites.
Personalizing reinforcement learning from human feedback with variational preference learning
Sriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta, and Natasha Jaques · 2024
Later among the works it cites.
Position: a roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Later among the works it cites.
Learning populations of preferences via pairwise comparison queries
Gokcan Tatli, Yi Chen, and Ramya Korlakai Vinayak · 2024
Later among the works it cites.
Metric learning from limited pairwise preference comparisons
Zhi Wang, Geelon So, and Ramya Korlakai Vinayak · 2024
Later among the works it cites.
Self-play preference optimization for language model alignment
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, and Quanquan Gu · 2024
Later among the works it cites.
Provable multi-party reinforcement learning with diverse human feedback
Huiying Zhong, Zhun Deng, Weijie J Su, Zhiwei Steven Wu, and Linjun Zhang · 2024
Later among the works it cites.
Metric distortion under probabilistic voting
Mohak Goyal and Sahasrajit Sarmasarkar · 2025
Closest in time.
Jackpot! alignment as a maximal lottery
Roberto-Rafael Maura-Rivero, Marc Lanctot, Francesco Visin, and Kate Larson · 2025
Closest in time.
Ariel D Procaccia, Benjamin Schiffer, and Shirley Zhang · 2025
Closest in time.
Direct alignment with heterogeneous preferences
Ali Shirali, Arash Nasr-Esfahany, Abdullah Alomar, Parsa Mirtaheri, Rediet Abebe, and Ariel Procaccia · 2025
Closest in time.