Fetching the paper…
Reading the bibliography…
Aligning generative models with human preference via RLHF typically suffers from overoptimization, where an imperfectly learned reward model can misguide the generative model to output undesired responses.
Fine-tuning language models from human preferences
Ziegler, D. M · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A · 1952
Earlier work this paper cites.
Minimax theorems
Fan, K · 1953
Earlier work this paper cites.
The covering number in learning theory
Zhou, D.-X · 2002
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Engstrom, L · 2005
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S · 2009
Earlier work this paper cites.
Understanding learned reward functions
Michaud, E. J · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y · 2012
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P · 2018
Earlier work this paper cites.
Program synthesis with large language models
Austin, J · 2021
Earlier work this paper cites.
Preference-based online learning with dueling bandits: A survey
Bengs, V · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Jin, Y · 2021
Earlier work this paper cites.
Dueling rl: reinforcement learning with trajectory preferences
Pacchiano, A · 2021
Earlier work this paper cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P · 2021
Earlier work this paper cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M · 2021
Earlier work this paper cites.
Towards instance-optimal offline reinforcement learning with pessimism
Yin, M · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y · 2022
Earlier work this paper cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X · 2022
Earlier work this paper cites.
Offline reinforcement learning with instrumental variables in confounded markov decision processes
Fu, Z · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D · 2022
Earlier work this paper cites.
Lu, M · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Cited alongside, same era.
Optimal conservative offline rl with general function approximation via augmented lagrangian
Rashidinejad, P · 2022
Cited alongside, same era.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L · 2022
Cited alongside, same era.
Causal confusion and reward misidentification in preference-based reward learning
Tien, J · 2022
Cited alongside, same era.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B · 2023
Later among the works it cites.
argilla-dpo-mix-7k
argill · 2024
Closest in time.
Double pessimism is provably efficient for distributionally robust offline reinforcement learning: Generic algorithm and robust partial coverage
Blanchet, J · 2024
Closest in time.
Exploration-driven policy optimization in rlhf: Theoretical insights on efficient data utilization
Du, Y · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiong, W · 2022
Cited alongside, same era.
Offline reinforcement learning with realizability and single-policy concentrability
Zhan, W · 2022
Cited alongside, same era.
Achiam, J · 2023
Cited alongside, same era.
Introducing claude
Anthropic · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S · 2023
Cited alongside, same era.
Reward model ensembles help mitigate overoptimization
Coste, T · 2023
Cited alongside, same era.
Dubois, Y · 2024
Closest in time.
Orpo: Monolithic preference optimization without reference model
Hong, J · 2024
Closest in time.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J · 2024
Closest in time.
Towards efficient and exact optimization of language model alignment
Ji, H · 2024
Closest in time.
Robust preference optimization with provable noise tolerance for llms
Liang, X · 2024
Closest in time.
Maximize to explore: One objective function fusing estimation, planning, and exploration
Liu, Z · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Pal, A · 2024
Closest in time.
From r r to q ⋆ q^{\star} : Your language model is secretly a q-function
Rafailov, R · 2024
Closest in time.
Countering reward over-optimization in llm with demonstration-guided reinforcement learning
Rita, M · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C · 2024
Closest in time.
Distributionally robust model-based offline reinforcement learning with near-optimal sample complexity
Shi, L · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y · 2024
Closest in time.
Is dpo superior to ppo for llm alignment? a comprehensive study
Xu, S · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Ye, C · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L · 2024
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H · 2024
Closest in time.
Iterative data smoothing: Mitigating reward overfitting and overoptimization in rlhf
Zhu, B · 2024
Closest in time.