Fetching the paper…
Reading the bibliography…
Policy alignment of large language models refers to constrained policy optimization, where the policy is optimized to maximize a reward while staying close to a reference policy with respect to an $f$-divergence such as the $\mathsf{KL}$ divergence.
On the theory of order statistics
A. Rényi · 1953
Earlier work this paper cites.
On rényi measures and hypothesis testing
O. Shayevitz · 2011
Earlier work this paper cites.
Concentration inequalities for order statistics
S. Boucheron and M. Thomas · 2012
Earlier work this paper cites.
Concentration Inequalities - A Nonasymptotic Theory of Independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Earlier work this paper cites.
Rényi divergence and kullback-leibler divergence
T. van Erven and P. Harremos · 2014
Earlier work this paper cites.
Bounds on the expectation of the maximum of samples from a gaussian
G. Kamath · 2015
Earlier work this paper cites.
Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling
F. Santambrogio · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
A variational characterization of rényi divergences
V. Anantharam · 2018
Earlier work this paper cites.
Interpolating between optimal transport and mmd using sinkhorn divergences, 2018
J. Feydy, T. Séjourné, F.-X. Vialard, S. ichi Amari, A. Trouvé, and G. Peyré · 2018
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Variational representations and neural network estimation of rényi divergences
J. Birrell, P. Dupuis, M. A. Katsoulakis, L. Rey-Bellet, and J. Wang · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
FUDGE: Controlled text generation with future discriminators
K. Yang and D. Klein · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Calibrating sequence likelihood improves conditional language generation
Y. Zhao, M. Khalman, R. Joshi, S. Narayan, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Theoretical guarantees on the best-of-n alignment policy, 2024
A. Beirami, A. Agarwal, J. Berant, A. D’Amour, J. Eisenstein, C. Nagpal, and A. T. Suresh · 2024
Closest in time.
Reward model ensembles help mitigate overoptimization
T. Coste, U. Anwar, R. Kirk, and D. Krueger · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring goodhart’s law, 2022
J. Hilton and L. Gao · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
Controlled decoding from language models
S. Mudgal, J. Lee, H. Ganapathy, Y. Li, T. Wang, Y. Huang, Z. Chen, H.-T. Cheng, M. Collins, T. Strohman, et al · 2023
Cited alongside, same era.
Information theory: From coding to learning, 2023
Y. Polyanskiy and Y. Wu · 2023
Cited alongside, same era.
Compositional preference models for aligning LMs
D. Go, T. Korbak, G. Kruszewski, J. Rozen, and M. Dymetman · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling, 2024
N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, N. A. Smith, and H. Hajishirzi · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Beyond reverse KL: Generalizing direct preference optimization with diverse divergence constraints
C. Wang, Y. Jiang, C. Yang, H. Liu, and Y. Chen · 2024
Closest in time.
Asymptotics of language model alignment, 2024
J. Q. Yang, S. Salamatian, Z. Sun, A. T. Suresh, and A. Beirami · 2024
Closest in time.