Fetching the paper…
Reading the bibliography…
The extraordinary capabilities of large language models (LLMs) such as ChatGPT and GPT-4 are in part unleashed by aligning them with reward models that are trained on human preferences, which are often represented as rankings of responses to prompts.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Foundations of modern potential theory
N. S. Landkof and N. Landkof · 1972
Earlier work this paper cites.
Crystalline order on a sphere and the generalized thomson problem
M. Bowick, A. Cacciuto, D. R. Nelson, and A. Travesset · 2002
Earlier work this paper cites.
Convex optimization
S. P. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Discretizing manifolds via minimum energy points
D. P. Hardin, E. B. Saff, et al · 2004
Earlier work this paper cites.
Asymptotics for minimal discrete riesz energy on curves in ℝ d \mathbb{R}^{d}
A. Martinez-Finkelshtein, V. Maymeskul, E. Rakhmanov, and E. Saff · 2004
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Expected absolute difference between two iid variables
S. L. (https://math.stackexchange.com/users/9340/sangchul lee) · 2017
Earlier work this paper cites.
Learning to understand goal specifications by modelling reward
D. Bahdanau, F. Hill, J. Leike, E. Hughes, A. Hosseini, P. Kohli, and E. Grefenstette · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Cited alongside, same era.
Thomson problem in one dimension: Minimal energy configurations of n charges on a curve
P. Amore and M. Jacobo · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Cited alongside, same era.
Axioms for learning from pairwise comparisons
R. Noothigattu, D. Peters, and A. D. Procaccia · 2020
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
V. Papyan, X. Han, and D. L. Donoho · 2020
Cited alongside, same era.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. H. Dodds, N. DasSarma, E. Tran-Johnson, et al · 2022
Later among the works it cites.
Teaching models to express their uncertainty in words
S. Lin, J. Hilton, and O. Evans · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Stanford human preferences dataset, 2023
K. Ethayarajh, H. Zhang, Y. Wang, and D. Jurafsky · 2023
Closest in time.
Longform: Optimizing instruction tuning for long text generation with corpus extraction, 2023
A. Köksal, T. Schick, A. Korhonen, and H. Schütze · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. He, J. Gao, and W. Chen · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Cited alongside, same era.
H. Liu, C. Sferrazza, and P. Abbeel · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Z. Sun, Y. Shen, Q. Zhou, H. Zhang, Z. Chen, D. Cox, Y. Yang, and C. Gan · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
B. Zhu, J. Jiao, and M. I. Jordan · 2023
Closest in time.