Fetching the paper…
Reading the bibliography…
Language model alignment methods such as reinforcement learning from human feedback (RLHF) have led to impressive advances in language model capabilities, but are limited by a widely observed phenomenon known as overoptimization, where the quality of the language model degrades over the course of the alignment process.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Aggregation of preference orderings
Germain Kreweras · 1965
Earlier work this paper cites.
On defining areas of voter choice: Professor tullock on stable voting
Paul B Simpson · 1969
Earlier work this paper cites.
On a class of equilibrium conditions for majority rule
Gerald H Kramer · 1973
Earlier work this paper cites.
Probabilistic social choice based on simple voting comparisons
Peter C Fishburn · 1984
Earlier work this paper cites.
Probability inequalities for likelihood ratios and convergence rates of sieve mles
Wing Hung Wong and Xiaotong Shen · 1995
Earlier work this paper cites.
On the Lambert W function
Robert M Corless, Gaston H Gonnet, David EG Hare, David J Jeffrey, and Donald E Knuth · 1996
Earlier work this paper cites.
Empirical Processes in M-Estimation
Sara A. Van de Geer · 2000
Earlier work this paper cites.
From ϵ \epsilon -entropy to KL-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B Tsybakov · 2008
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Rényi divergence and kullback-leibler divergence
Tim Van Erven and Peter Harremos · 2014
Earlier work this paper cites.
Contextual dueling bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
Nihar Shah, Sivaraman Balakrishnan, Joseph Bradley, Abhay Parekh, Kannan Ramchandran, and Martin Wainwright · 2015
Earlier work this paper cites.
Boltzmann exploration done right
Nicolò Cesa-Bianchi, Claudio Gentile, Gábor Lugosi, and Gergely Neu · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Semi-parametric efficient policy learning with continuous actions
Victor Chernozhukov, Mert Demirer, Greg Lewis, and Vasilis Syrgkanis · 2019
Earlier work this paper cites.
Variance-based regularization with convex objectives
John Duchi and Hongseok Namkoong · 2019
Earlier work this paper cites.
FLAMBE: Structural complexity and representation learning of low rank MDPs
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Earlier work this paper cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Nathan Kallus and Masatoshi Uehara · 2020
Earlier work this paper cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Earlier work this paper cites.
Provably good batch off-policy reinforcement learning without great exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2020
Earlier work this paper cites.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2020
Earlier work this paper cites.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, and Shengyi Huang · 2020
Earlier work this paper cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Earlier work this paper cites.
Preference-based reinforcement learning with finite-time guarantees
Yichong Xu, Ruosong Wang, Lin Yang, Aarti Singh, and Artur Dubrawski · 2020
Earlier work this paper cites.
Off-policy imitation learning from observations
Zhuangdi Zhu, Kaixiang Lin, Bo Dai, and Jiayu Zhou · 2020
Earlier work this paper cites.
Policy learning with observational data
Susan Athey and Stefan Wager · 2021
Earlier work this paper cites.
Is pessimism provably efficient for offline RL?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Cited alongside, same era.
Optidice: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byungjun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Cited alongside, same era.
Dueling RL: Reinforcement learning with trajectory preferences
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Cited alongside, same era.
Gibbs sampling from human feedback: A provable KL-constrained framework for RLHF
Wei Xiong, Hanze Dong, Chenlu Ye, Han Zhong, Nan Jiang, and Tong Zhang · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Later among the works it cites.
Scalable online exploration via coverability
Philip Amortila, Dylan J Foster, and Akshay Krishnamurthy · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Closest in time.
Value-incentivized preference optimization: A unified approach to online and offline RLHF
Shicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai, Tong Yang, Sherry Yang, Dale Schuurmans, Yuejie Chi, and Bo Dai · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Cited alongside, same era.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, and Sham M Kakade · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Offline reinforcement learning under value and density-ratio realizability: The power of gaps
Jinglin Chen and Nan Jiang · 2022
Cited alongside, same era.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Cited alongside, same era.
When are offline two-player zero-sum Markov games solvable?
Qiwen Cui and Simon S Du · 2022
Cited alongside, same era.
Closest in time.
Dataset reset policy optimization for RLHF
Jonathan D Chang, Wenhao Shan, Owen Oertell, Kianté Brantley, Dipendra Misra, Jason D Lee, and Wen Sun · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Closest in time.
Provably sample efficient RLHF via active preference optimization
Nirjhar Das, Souradip Chakraborty, Aldo Pacchiano, and Sayak Ray Chowdhury · 2024
Closest in time.
RLHF workflow: From reward modeling to online RLHF
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang · 2024
Closest in time.
Exploration-driven policy optimization in RLHF: Theoretical insights on efficient data utilization
Yihan Du, Anna Winnicki, Gal Dalal, Shie Mannor, and R Srikant · 2024
Closest in time.
Robust preference optimization through reward model distillation
Adam Fisch, Jacob Eisenstein, Vicky Zayats, Alekh Agarwal, Ahmad Beirami, Chirag Nagpal, Pete Shaw, and Jonathan Berant · 2024
Closest in time.
Importance-weighted offline learning done right
Germano Gabbianelli, Gergely Neu, and Matteo Papini · 2024
Closest in time.
REBEL: Reinforcement learning via regressing relative rewards
Zhaolin Gao, Jonathan D Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, J Andrew Bagnell, Jason D Lee, and Wen Sun · 2024
Closest in time.
Direct language model alignment from online AI feedback
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao, Bilal Piot, Johan Ferret, and Mathieu Blondel · 2024
Closest in time.
Self-play with adversarial critic: Provable and scalable offline alignment for language models
Xiang Ji, Sanjeev Kulkarni, Mengdi Wang, and Tengyang Xie · 2024
Closest in time.
Provably mitigating overoptimization in RLHF: Your SFT loss is implicitly an adversarial regularizer
Zhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu, Hongyi Guo, Yingxiang Yang, Jose Blanchet, and Zhaoran Wang · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with DPO-positive
Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White · 2024
Closest in time.
Countering reward over-optimization in LLM with demonstration-guided reinforcement learning
Mathieu Rita, Florian Strub, Rahma Chaabouni, Paul Michel, Emmanuel Dupoux, and Olivier Pietquin · 2024
Closest in time.
Direct Nash Optimization: Teaching language models to self-improve with general preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie · 2024
Closest in time.
Understanding preference fine-tuning through the lens of coverage
Yuda Song, Gokul Swamy, Aarti Singh, J Andrew Bagnell, and Wen Sun · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Closest in time.
Preference fine-tuning of LLMs should leverage suboptimal, on-policy data
Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn, and Aviral Kumar · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot · 2024
Closest in time.
Oracle-efficient pessimism: Offline policy optimization in contextual bandits
Lequn Wang, Akshay Krishnamurthy, and Alex Slivkins · 2024
Closest in time.
Self-play preference optimization for language model alignment
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, and Quanquan Gu · 2024
Closest in time.
Exploratory preference optimization: Harnessing implicit Q*-approximation for sample-efficient rlhf
Tengyang Xie, Dylan J Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah, and Alexander Rakhlin · 2024
Closest in time.
A theoretical analysis of Nash learning from human feedback under general KL-regularized preference
Chenlu Ye, Wei Xiong, Yuheng Zhang, Nan Jiang, and Tong Zhang · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Xiaoying Zhang, Jean-Francois Ton, Wei Shen, Hongning Wang, and Yang Liu · 2024
Closest in time.
Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF
Banghua Zhu, Michael I Jordan, and Jiantao Jiao · 2024
Closest in time.
Provably efficient offline goal-conditioned reinforcement learning with general function approximation and single-policy concentrability
Hanlin Zhu and Amy Zhang · 2024
Closest in time.