Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has emerged as a reliable approach to aligning large language models (LLMs) to human preferences.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
TAMER: Training an agent manually via evaluative reinforcement
W. B. Knox and P. Stone · 2008
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
OpenAI baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
R. Paulus, C. Xiong, and R. Socher · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Learning dynamic robot-to-human object handover from human feedback
A. Kupcsik, D. Hsu, and W. S. Lee · 2018
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Personalized reward learning with interaction-grounded learning (IGL)
J. Maghakian, P. Mineiro, K. Panaganti, M. Rucker, A. Saran, and C. Tan · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2022
Later among the works it cites.
Offline RL for natural language generation with implicit language Q-learning
C. Snell, I. Kostrikov, Y. Su, M. Yang, and S. Levine · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
A general language assistant as a laboratory for alignment, 2021
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, J. Kernion, K. Ndousse, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, and J. Kaplan · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
Simple agent, complex environment: Efficient reinforcement learning with agent states
S. Dong, B. Van Roy, and Z. Zhou · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al
Cited in the paper.
A general sample complexity analysis of vanilla policy gradient
R. Yuan, R. M. Gower, and A. Lazaric · 2022
Later among the works it cites.
StackLLaMA: An RL fine-tuned LLaMA model for Stack Exchange question and answering, 2023
E. Beeching, Y. Belkada, K. Rasul, L. Tunstall, L. von Werra, N. Rajani, and N. Lambert · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
B. Zhu, J. Jiao, and M. I. Jordan · 2023
Closest in time.