Fetching the paper…
Reading the bibliography…
A common approach for aligning language models to human preferences is to first learn a reward model from preference data, and then use this reward model to update the language model.
“Rank analysis of incomplete block designs: I. The method of paired comparisons”
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel et al · 2020
Earlier work this paper cites.
“Learning to summarize with human feedback”
Nisan Stiennon et al · 2020
Earlier work this paper cites.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
“Fine-tuning language models to find agreement among humans with diverse preferences”, 2022
Michiel. Bakker et al · 2022
Earlier work this paper cites.
“Nano: Nested Human-in-the-Loop Reward Learning for Few-shot Language Model Control”
Xiang Fan et al · 2022
Earlier work this paper cites.
“RL with KL penalties is better viewed as Bayesian inference”
Tomasz Korbak, Ethan Perez and Christopher Buckley · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Cited alongside, same era.
“Calibrating sequence likelihood improves conditional language generation”
Yao Zhao et al · 2022
Cited alongside, same era.
Rohan Anil et al · 2023
Cited alongside, same era.
“A general theoretical paradigm to understand learning from human preferences”
Mohammad Azar et al · 2023
Cited alongside, same era.
“Reward model ensembles help mitigate overoptimization”
Thomas Coste, Usman Anwar, Robert Kirk and David Krueger · 2023
Cited alongside, same era.
“Large language models are zero-shot rankers for recommender systems”
Yupeng Hou et al · 2023
Later among the works it cites.
“Confronting Reward Model Overoptimization with Constrained RLHF”, 2023
Ted Moskovitz et al · 2023
Later among the works it cites.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2023
Later among the works it cites.
“The trickle-down impact of reward (in-) consistency on rlhf”
Lingfeng Shen et al · 2023
Later among the works it cites.
“A long way to go: Investigating length correlations in rlhf”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Alpacafarm: A simulation framework for methods that learn from human feedback”
Yann Dubois et al · 2023
Cited alongside, same era.
“Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking”
Jacob Eisenstein et al · 2023
Cited alongside, same era.
“Scaling laws for reward model overoptimization”
Leo Gao, John Schulman and Jacob Hilton · 2023
Cited alongside, same era.
Prasann Singhal, Tanya Goyal, Jiacheng Xu and Greg Durrett · 2023
Later among the works it cites.
“Fine-Grained Human Feedback Gives Better Rewards for Language Model Training”, 2023
Zeqiu Wu et al · 2023
Later among the works it cites.
Yuanzhao Zhai et al · 2023
Later among the works it cites.
“WARM: On the Benefits of Weight Averaged Reward Models”
Alexandre Ram\’e et al · 2024
Closest in time.