Fetching the paper…
Reading the bibliography…
This paper concerns the problem of aligning samples from large language models to human preferences using best-of-$n$ sampling, where we draw $n$ samples, rank them, and return the best one.
“Rank analysis of incomplete block designs: I. The method of paired comparisons”
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
“Lectures on modern convex optimization: analysis, algorithms, and engineering applications”
Aharon Ben-Tal and Arkadi Nemirovski · 2001
Earlier work this paper cites.
“Deep Reinforcement Learning from Human Preferences”
Paul Christiano et al · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman et al · 2017
Earlier work this paper cites.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“Learning to summarize with human feedback”
Nisan Stiennon et al · 2020
Earlier work this paper cites.
“Webgpt: Browser-assisted question-answering with human feedback”
Reiichiro Nakano et al · 2021
Earlier work this paper cites.
“FUDGE: Controlled Text Generation With Future Discriminators”
Kevin Yang and Dan Klein · 2021
Earlier work this paper cites.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
“Measuring Goodhart’s law”, 2022
Leo Jacob · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“COLD Decoding: Energy-based Constrained Text Generation with Langevin Dynamics”
Lianhui Qin, Sean Welleck, Daniel Khashabi and Yejin Choi · 2022
Earlier work this paper cites.
“Defining and characterizing reward gaming”
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov and David Krueger · 2022
Earlier work this paper cites.
“Pythia: A suite for analyzing large language models across training and scaling”
Stella Biderman et al · 2023
Earlier work this paper cites.
“RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment”
Hanze Dong et al · 2023
Earlier work this paper cites.
“Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking”
Jacob Eisenstein et al · 2023
Cited alongside, same era.
“Scaling laws for reward model overoptimization”
Leo Gao, John Schulman and Jacob Hilton · 2023
Cited alongside, same era.
“Reinforced self-training (rest) for language modeling”
Caglar Gulcehre et al · 2023
Cited alongside, same era.
“A survey of reinforcement learning from human feedback”
Timo Kaufmann, Paul Weng, Viktor Bengs and Eyke H\"ullermeier · 2023
Cited alongside, same era.
“Statistical rejection sampling improves preference optimization”
Tianqi Liu et al · 2023
Cited alongside, same era.
“Secrets of rlhf in large language models part i: Ppo”
Rui Zheng et al · 2023
Later among the works it cites.
“A general theoretical paradigm to understand learning from human preferences”
Mohammad Azar et al · 2024
Closest in time.
“Theoretical guarantees on the best-of-n alignment policy”
Ahmad Beirami et al · 2024
Closest in time.
“Kto: Model alignment as prospect theoretic optimization”
Kawin Ethayarajh et al · 2024
Closest in time.
“Reference-free monolithic preference optimization with odds ratio”
Jiwoo Hong, Noah Lee and James Thorne · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Controlled decoding from language models”
Sidharth Mudgal et al · 2023
Cited alongside, same era.
“Direct Preference Optimization: Your Language Model is Secretly a Reward Model”
Rafael Rafailov et al · 2023
Cited alongside, same era.
“Efficient rlhf: Reducing the memory usage of ppo”
Michael Santacroce et al · 2023
Cited alongside, same era.
“Large language model alignment: A survey”
Tianhao Shen et al · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“Aligning large language models with human: A survey”
Yufei Wang et al · 2023
Cited alongside, same era.
“Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint”
Wei Xiong et al · 2023
Cited alongside, same era.
“Regularized Best-of-N Sampling to Mitigate Reward Hacking for Language Model Alignment”
Yuu Jinnai, Tetsuro Morimura, Kaito Ariu and Kenshi Abe · 2024
Closest in time.
“Preventing reward hacking with occupancy measure regularization”
Cassidy Laidlaw, Shivam Singhal and Anca Dragan · 2024
Closest in time.
“Inference-time intervention: Eliciting truthful answers from a language model”
Kenneth Li et al · 2024
Closest in time.
“Statistical Rejection Sampling Improves Preference Optimization”
Tianqi Liu et al · 2024
Closest in time.
“Disentangling length from quality in direct preference optimization”
Ryan Park, Rafael Rafailov, Stefano Ermon and Chelsea Finn · 2024
Closest in time.
“Understanding the performance gap between online and offline alignment algorithms”, 2024
Yunhao Tang et al · 2024
Closest in time.
“Transforming and Combining Rewards for Aligning Large Language Models”, 2024
Zihao Wang et al · 2024
Closest in time.
Haoran Xu et al · 2024
Closest in time.
“Asymptotics of Language Model Alignment”
Joy Yang et al · 2024
Closest in time.