Fetching the paper…
Reading the bibliography…
Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assuming, implicitly or explicitly, a model of human preferences.
Fisher’s exact test
Graham JG Upton · 1992
Earlier work this paper cites.
Thinking, Fast and Slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca Dragan, Shankar Sastry, and Sanjit Seshia · 2017
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
Erdem Bıyık, Dylan P Losey, Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
Erdem Bıyık, Dylan P Losey, Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Earlier work this paper cites.
Models of human preference for learning reward functions
W Bradley Knox, Stephane Hatgis-Kessell, Serena Booth, Scott Niekum, Peter Stone, and Alessandro Allievi · 2022
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Skill preferences: Learning to extract and execute robotic skills from human feedback
Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi, Pieter Abbeel, and Michael Laskin · 2022
Cited alongside, same era.
Moral machine or tyranny of the majority?
Michael Feffer, Hoda Heidari, and Zachary C Lipton · 2023
Cited alongside, same era.
Inverse constitutional ai: Compressing preferences into principles
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier, Samuel Albanie, and Robert Mullins · 2024
Later among the works it cites.
Learning optimal advantage from preferences and mistaking it for reward
W Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson, Serena Booth, Anca Dragan, Peter Stone, and Scott Niekum · 2024
Later among the works it cites.
Choice between partial trajectories: Disentangling goals from beliefs
Henrik Marklund and Benjamin Van Roy · 2024
Later among the works it cites.
Value imprint: A technique for auditing the human values embedded in rlhf datasets
Ike Obi, Rohan Pant, Srishti Shekhar Agrawal, Maham Ghazanfar, and Aaron Basiletti · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn, Scott Niekum, W Bradley Knox, and Dorsa Sadigh · 2023
Cited alongside, same era.
Few-shot preference learning for human-in-the-loop rl
Donald Joseph Hejna III and Dorsa Sadigh · 2023
Cited alongside, same era.
Preference transformer: Modeling human preferences using transformers for rl
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2023
Cited alongside, same era.
RLAIF: Scaling reinforcement learning from human feedback with AI feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Cited alongside, same era.
Kimin Lee, Laura Smith, and Pieter Abbeel
Cited in the paper.
Later among the works it cites.
The consequences of ai training on human decision-making
Lauren S Treiman, Chien-Ju Ho, and Wouter Kool · 2024
Later among the works it cites.
Decoding global preferences: Temporal and cooperative dependency modeling in multi-agent preference-based reinforcement learning
Tianchen Zhu, Yue Qiu, Haoyi Zhou, and Jianxin Li · 2024
Later among the works it cites.
Ai alignment at your discretion
Maarten Buyl, Hadi Khalaf, Claudio Mayrink Verdun, Lucas Monteiro Paes, Caio Cesar Vieira Machado, and Flavio du Pin Calmon · 2025
Closest in time.
Do people think fast or slow when training ai?
Lauren S. Treiman, Chien-Ju Ho, and Wouter Kool · 2025
Closest in time.
Helpful, harmless, honest? rlhf as survey design and content moderation
Samantha D’Alonzo, Frauke Kreuter, and Serena Booth · 2026
Closest in time.