Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models.
An invariant form for the prior probability in estimation problems
H. Jeffreys · 1946
Earlier work this paper cites.
Information Theory and Statistics
S. Kullback · 1959
Earlier work this paper cites.
Semi-distributed representations and catastrophic forgetting in connectionist networks
R. M. French · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Elements of information theory
T. M. Cover · 1999
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Faulty Reward Functions in the Wild
J. Clark and D. Amodei · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Earlier work this paper cites.
Iterated distillation and amplification
A. Cotra · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Earlier work this paper cites.
Buy 4 REINFORCE samples, get a baseline for free!
W. Kool, H. van Hoof, and M. Welling · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Multi-agent communication meets natural language: Synergies between functional and structural language learning
A. Lazaridou, A. Potapenko, and O. Tieleman · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Leverage the average: an analysis of kl regularization in reinforcement learning
N. Vieillard, T. Kozuno, B. Scherrer, O. Pietquin, R. Munos, and M. Geist · 2020
Cited alongside, same era.
A general language assistant as a laboratory for alignment
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, J. Kernion, K. Ndousse, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, and J. Kaplan · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
On-policy distillation of language models: Learning from self-generated mistakes
R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem · 2024
Closest in time.
Back to basics: Revisiting REINFORCE style optimization for learning from human feedback in LLMs
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, A. Üstün, and S. Hooker · 2024
Closest in time.
Variational Best-of-N alignment
A. Amini, T. Vieira, and R. Cotterell · 2024
Closest in time.
Theoretical guarantees on the Best-of-N alignment policy
A. Beirami, A. Agarwal, J. Berant, A. D’Amour, J. Eisenstein, C. Nagpal, and A. T. Suresh · 2024
Closest in time.
Recurrentgemma: Moving past transformers for efficient open language models
A. Botev, S. De, S. L. Smith, A. Fernando, G.-C. Muraru, R. Haroun, L. Berrada, R. Pascanu, P. G. Sessa, R. Dadashi, L. Hussenot, J. Ferret, S. Girgin, O. Bachem, A. Andreev, K. Kenealy, T. Mesnard, C. Hardin, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, P. Tafti, A. Joulin, N. Fiedel, E. Senter, Y. Chen, S. Srinivasan, G. Desjardins, D. Budden, A. Doucet, S. Vikram, A. Paszke, T. Gale, S. Borgeaud, C. Chen, A. Brock, A. Paterson, J. Brennan, M. Risdal, R. Gundluru, N. Devanathan, P. Mooney, N. Chauhan, P. Culliton, L. G. Martins, E. Bandy, D. Huntsperger, G. Cameron, A. Zucker, T. Warkentin, L. Peran, M. Giang, Z. Ghahramani, C. Farabet, K. Kavukcuoglu, D. Hassabis, R. Hadsell, Y. W. Teh, and N. de Frietas · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
A. Pan, K. Bhatia, and J. Steinhardt · 2022
Cited alongside, same era.
Defining and characterizing reward gaming
J. M. V. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2022
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Cited alongside, same era.
RAFT: Reward ranked finetuning for generative foundation model alignment
H. Dong, W. Xiong, D. Goyal, Y. Zhang, W. Chow, R. Pan, S. Diao, J. Zhang, K. SHUM, and T. Zhang · 2023
Cited alongside, same era.
Closest in time.
Human alignment of large language models through online preference optimisation
D. Calandriello, D. Guo, R. Munos, M. Rowland, Y. Tang, B. A. Pires, P. H. Richemond, C. L. Lan, M. Valko, T. Liu, et al · 2024
Closest in time.
Codegemma: Open code models based on gemma
CodeGemma Team · 2024
Closest in time.
Contrastive policy gradient: Aligning llms on sequence-level scores in a supervised-friendly fashion
Y. Flet-Berliac, N. Grinsztajn, F. Strub, E. Choi, C. Cremer, A. Ahmadian, Y. Chandak, M. G. Azar, O. Pietquin, and M. Geist · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team · 2024
Closest in time.
Learn your reference model for real good alignment
A. Gorbatovski, B. Shaposhnikov, A. Malakhov, N. Surnachev, Y. Aksenov, I. Maksimov, N. Balagansky, and D. Gavrilov · 2024
Closest in time.
Bonbon alignment for large language models and the sweetness of Best-of-N sampling
L. Gui, C. Gârbacea, and V. Veitch · 2024
Closest in time.
Direct language model alignment from online ai feedback
S. Guo, B. Zhang, T. Liu, T. Liu, M. Khalman, F. Llinares, A. Rame, T. Mesnard, Y. Zhao, B. Piot, J. Ferret, and M. Blondel · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
West-of-n: Synthetic preference generation for improved reward modeling
A. Pace, J. Mallinson, E. Malmi, S. Krause, and A. Severyn · 2024
Closest in time.
Asymptotics of language model alignment
J. Qiping Yang, S. Salamatian, Z. Sun, A. Theertha Suresh, and A. Beirami · 2024
Closest in time.
WARM: On the benefits of weight averaged reward models
A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret · 2024
Closest in time.
WARP: On the benefits of weight averaged rewarded policies
A. Ramé, J. Ferret, N. Vieillard, R. Dadashi, L. Hussenot, P.-L. Cedoz, P. G. Sessa, S. Girgin, A. Douillard, and O. Bachem · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms
Y. Tang, D. Z. Guo, Z. Zheng, D. Calandriello, Y. Cao, E. Tarassov, R. Munos, B. Ávila Pires, M. Valko, Y. Cheng, and W. Dabney · 2024
Closest in time.
Self-play preference optimization for language model alignment
Y. Wu, Z. Sun, H. Yuan, K. Ji, Y. Yang, and Q. Gu · 2024
Closest in time.
Asymptotics of language model alignment
J. Q. Yang, S. Salamatian, Z. Sun, A. T. Suresh, and A. Beirami · 2024
Closest in time.