Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has become essential for improving language model capabilities, but traditional approaches rely on the assumption that human preferences follow a transitive Bradley-Terry model.
Zur theorie der gesellschaftsspiele
John von Neumann · 1928
Earlier work this paper cites.
Equilibrium points in n-person games
John Nash · 1950
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
Intransitivity, utility, and the aggregation of preference patterns
Kenneth O May · 1954
Earlier work this paper cites.
Quantal response equilibria for normal form games
Richard D. McKelvey and Thomas R. Palfrey · 1995
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire · 1999
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Constantinos Daskalakis, Alan Deckelbaum, and Anthony Kim · 2011
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: A preliminary survey
Christian Wirth and Johannes Fürnkranz · 2013
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Francis Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Earlier work this paper cites.
Multiplicative weights update in zero-sum games
James P Bailey and Georgios Piliouras · 2018
Earlier work this paper cites.
Last-iterate convergence: Zero-sum games and constrained min-max optimization
Constantinos Daskalakis and Ioannis Panageas · 2018
Earlier work this paper cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec · 2020
Earlier work this paper cites.
Linear last-iterate convergence in constrained saddle-point optimization
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, and Haipeng Luo · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeff Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Earlier work this paper cites.
Fast policy extragradient methods for competitive games with entropy regularization
Shicong Cen, Yuting Wei, and Yuejie Chi · 2021
Cited alongside, same era.
Advances in preference-based reinforcement learning: A review
Youssef Abdelkareem, Shady Shehata, and Fakhri Karray · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Anthropic · 2022
Cited alongside, same era.
Faster last-iterate convergence of policy optimization in zero-sum markov games
Shicong Cen, Yuejie Chi, Simon S Du, and Lin Xiao · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Cited alongside, same era.
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen · 2024
Later among the works it cites.
Lecture Notes for High-Dimensional Statistics - 18.S997, Spring 2015
Philippe Rigollet · 2024
Later among the works it cites.
From r r to q ∗ q^{*} : Your language model is secretly a q-function, 2024
Rafael Rafailov, Joey Hejna, Ryan Park, and Chelsea Finn · 2024
Later among the works it cites.
Direct nash optimization: Teaching language models to self-improve with general preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie · 2024
Later among the works it cites.
Multi-turn reinforcement learning from preference human feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peft: State-of-the-art parameter-efficient fine-tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Samuel Sokota, Ryan D’Orazio, J Zico Kolter, Nicolas Loizou, Marc Lanctot, Ioannis Mitliagkas, Noam Brown, and Christian Kroer · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Cited alongside, same era.
Nash learning from human feedback
Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Andrea Michi, et al · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, Idan Szpektor, et al · 2024
Later among the works it cites.
The importance of online data: Understanding preference fine-tuning via coverage, 2024
Yuda Song, Gokul Swamy, Aarti Singh, J. Andrew Bagnell, and Wen Sun · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Later among the works it cites.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Fahim Tajwar, Anika Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn, and Aviral Kumar · 2024
Later among the works it cites.
Generalized preference optimization: A unified approach to offline alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot · 2024
Later among the works it cites.
Mingzhi Wang, Chengdong Ma, Qizhi Chen, Linjian Meng, Yang Han, Jiancong Xiao, Zhaowei Zhang, Jing Huo, Weijie J. Su, and Yaodong Yang · 2024
Later among the works it cites.
Self-play preference optimization for language model alignment
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, and Quanquan Gu · 2024
Later among the works it cites.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Tengyang Xie, Dylan J Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah, and Alexander Rakhlin · 2024
Later among the works it cites.
Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL-constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim · 2024
Later among the works it cites.
Online iterative reinforcement learning from human feedback with general preference model
Chenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Iterative nash policy optimization: Aligning llms with general preferences via no-regret learning
Yuheng Zhang, Dian Yu, Baolin Peng, Linfeng Song, Ye Tian, Mingyue Huo, Nan Jiang, Haitao Mi, and Dong Yu · 2024
Later among the works it cites.
Mingyu Chen, Yiding Chen, Wen Sun, and Xuezhou Zhang · 2025
Closest in time.
Pilaf: Optimal human preference sampling for reward modeling
Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng, Julia Kempe, and Yaqi Duan · 2025
Closest in time.
The crucial role of samplers in online direct preference optimization
Ruizhe Shi, Runlong Zhou, and Simon Shaolei Du · 2025
Closest in time.
Game-theoretic regularized self-play alignment of large language models
Xiaohang Tang, Sangwoong Yoon, Seongho Son, Huizhuo Yuan, Quanquan Gu, and Ilija Bogunovic · 2025
Closest in time.
Improving llm general preference alignment via optimistic online mirror descent
Yuheng Zhang, Dian Yu, Tao Ge, Linfeng Song, Zhichen Zeng, Haitao Mi, Nan Jiang, and Dong Yu · 2025
Closest in time.