Fetching the paper…
Reading the bibliography…
In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO).
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Tamer: Training an Agent Manually via Evaluative Reinforcement
Knox, W. B. and Stone, P · 2008
Earlier work this paper cites.
The K-armed Dueling Bandits Problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T · 2009
Earlier work this paper cites.
Reducing Dueling Bandits to Cardinal Bandits
Ailon, N., Karnin, Z. S., and Joachims, T · 2014
Earlier work this paper cites.
Relative Confidence Sampling for Efficient On-line Ranker Evaluation
Zoghi, M., Whiteson, S. A., De Rijke, M., and Munos, R · 2014
Earlier work this paper cites.
A Relative Exponential Weighing Algorithm for Adversarial Utility-based Dueling Bandits
Gajane, P., Urvoy, T., and Clérot, F · 2015
Earlier work this paper cites.
Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem
Komiyama, J., Honda, J., Kashima, H., and Nakagawa, H · 2015
Earlier work this paper cites.
Linear Convergence of Gradient and Proximal-Gradient Methods under the Polyak-Łojasiewicz Condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Earlier work this paper cites.
Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M. J., et al · 2016
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Interactive Learning from Policy-dependent Human Feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L · 2017
Earlier work this paper cites.
Deep TAMER: Interactive Agent Shaping in High-dimensional State Spaces
Warnell, G., Waytowich, N., Lawhern, V., and Stone, P · 2018
Earlier work this paper cites.
Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
Brown, D. S., Goo, W., Nagarajan, P., and Niekum, S · 2019
Earlier work this paper cites.
Off-policy Deep Reinforcement Learning without Exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Way Off-policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Earlier work this paper cites.
Safe Policy Improvement with Baseline Bootstrapping
Laroche, R., Trichelair, P., and Des Combes, R. T · 2019
Earlier work this paper cites.
Algaedice: Policy Gradient from Arbitrary Experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D · 2019
Earlier work this paper cites.
Active Ranking with Subset-wise Preferences
Saha, A. and Gopalan, A · 2019
Earlier work this paper cites.
Fine-tuning Language Models from Human Preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
An Optimistic Perspective on Offline Reinforcement Learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Improved Optimistic Algorithms for Logistic Bandits
Faury, L., Abeille, M., Calauzènes, C., and Fercoq, O · 2020
Cited alongside, same era.
Morel: Model-based Offline Rinforcement Learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative Q-learning for Offline Reinforcement Learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
On the Global Convergence Rates of Softmax Policy Gradient Methods
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Ganguli, D. et al · 2022
Later among the works it cites.
Improving Alignment of Dialogue Agents via Targeted Human Judgements
Glaese, A. et al · 2022
Later among the works it cites.
Teaching Language Models to Support Answers with Verified Quotes
Menick, J. et al · 2022
Later among the works it cites.
Training Language Models to Follow Instructions with Human Feedback
Ouyang, L. et al · 2022
Later among the works it cites.
Efficient and Optimal Algorithms for Contextual Dueling Bandits under Realizability
Saha, A. and Krishnamurthy, A · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mei, J., Xiao, C., Szepesvári, C., and Schuurmans, D · 2020
Cited alongside, same era.
Learning to Summarize with Human Feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Cited alongside, same era.
On the Theory of Reinforcement Learning with Once-per-episode Feedback
Chatterji, N., Pacchiano, A., Bartlett, P., and Jordan, M · 2021
Cited alongside, same era.
Is Pessimism Provably Efficient for Offline RL?
Jin, Y., Yang, Z., and Wang, Z · 2021
Cited alongside, same era.
Optidice: Offline Policy Optimization via Stationary Distribution Correction Estimation
Lee, J., Jeon, W., Lee, B., Pineau, J., and Kim, K.-E · 2021
Cited alongside, same era.
Webgpt: Browser-assisted Question-answering with Human Feedback
Nakano, R. et al · 2021
Cited alongside, same era.
Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of Pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Cited alongside, same era.
Direct Preference-based Policy Optimization without Reward Modeling
An, G., Lee, J., Zuo, X., Kosaka, N., Kim, K.-M., and Song, H. O · 2023
Later among the works it cites.
A General Theoretical Paradigm to Understand Learning from Human Preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Later among the works it cites.
Scaling Laws for Reward Model Overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Later among the works it cites.
Contrastive Prefence Learning: Learning from Human Feedback without RL
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Is Reinforcement Learning (not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2023
Later among the works it cites.
Dueling RL: Reinforcement Learning with Trajectory Preferences
Saha, A., Pacchiano, A., and Lee, J · 2023
Later among the works it cites.
Benchmarks and Algorithms for Offline Preference-Based Reward Learning
Shin, D., Dragan, A. D., and Brown, D. S · 2023
Later among the works it cites.
Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Wang, C. et al · 2023
Later among the works it cites.
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint
Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T · 2023
Later among the works it cites.
Provable Offline Reinforcement Learning with Human Feedback
Zhan, W., Uehara, M., Kallus, N., Lee, J. D., and Sun, W · 2023
Later among the works it cites.
Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
Zhu, B., Jordan, M. I., and Jiao, J · 2023
Later among the works it cites.