Fetching the paper…
Reading the bibliography…
Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning from human feedback (RLHF), which forces the learned policy to stay close to a reference policy.
Fine-tuning language models from human preferences
Ziegler, D. M · 1909
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J · 2007
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D · 2013
Earlier work this paper cites.
Trust region policy optimization
Schulman, J · 2015
Earlier work this paper cites.
Brockman, G · 2016
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R · 2016
Earlier work this paper cites.
Conservative bandits
Wu, Y · 2016
Earlier work this paper cites.
Fast rates for bandit optimization with upper-confidence frank-wolfe
Berthet, Q · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
Understanding the impact of entropy on policy optimization
Ahmed, Z · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Geist, M · 2019
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A · 2020
Earlier work this paper cites.
Bandit algorithms
Lattimore, T · 2020
Earlier work this paper cites.
On the global convergence rates of softmax policy gradient methods
Mei, J · 2020
Earlier work this paper cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L · 2020
Earlier work this paper cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Wang, R · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K · 2021
Cited alongside, same era.
The statistical complexity of interactive decision making
Foster, D. J · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Hendrycks, D · 2021
Cited alongside, same era.
Corruption-robust algorithms with uncertainty weighting for nonlinear contextual bandits and markov decision processes
Ye, C · 2023
Later among the works it cites.
Mathematical analysis of machine learning algorithms
Zhang, T · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B · 2023
Later among the works it cites.
The non-linear f f -design and applications to interactive learning
Agarwal, A · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2024
Closest in time.
Human alignment of large language models through online preference optimisation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Cited alongside, same era.
Algorithms for reinforcement learning
Szepesvári, C · 2022
Cited alongside, same era.
Achiam, J · 2023
Cited alongside, same era.
Vo q q l: Towards optimal regret in model-free rl with nonlinear function approximation
Agarwal, A · 2023
Cited alongside, same era.
Introducing claude
Anthropic, A · 2023
Cited alongside, same era.
Calandriello, D · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf
Dong, H · 2024
Closest in time.
ToRA: A tool-integrated reasoning agent for mathematical problem solving
Gou, Z · 2024
Closest in time.
Bonbon alignment for large language models and the sweetness of best-of-n sampling
Gui, L · 2024
Closest in time.
Direct language model alignment from online ai feedback
Guo, S · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta, A · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z · 2024
Closest in time.
The importance of online data: Understanding preference fine-tuning via coverage
Song, Y · 2024
Closest in time.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
Tong, Y · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y · 2024
Closest in time.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Xie, T · 2024
Closest in time.
From lists to emojis: How format bias affects model alignment
Zhang, X · 2024
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H · 2024
Closest in time.