Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) is a powerful technique for training agents to perform difficult-to-specify tasks.
Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 1904
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences, January 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning, December 2020
Hong Jun Jeon, Smitha Milli, and Anca D. Dragan · 2002
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2009
Earlier work this paper cites.
Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise
Jacob Whitehill, Ting-fan Wu, Jacob Bergsma, Javier Movellan, and Paul Ruvolo · 2009
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D Ziebart, J Andrew Bagnell, and Anind K Dey · 2010
Earlier work this paper cites.
Learning from Suboptimal Demonstration via Self-Supervised Reward Regression, November 2020
Letian Chen, Rohan Paleja, and Matthew Gombolay · 2010
Earlier work this paper cites.
Learning From Crowds
Vikas C. Raykar, Shipeng Yu, Linda H. Zhao, Gerardo Hermosillo Valadez, Charles Florin, Luca Bogoni, and Linda Moy · 2010
Earlier work this paper cites.
The Multidimensional Wisdom of Crowds
Peter Welinder, Steve Branson, Pietro Perona, and Serge Belongie · 2010
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning, June 2016
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Faulty reward functions in the wild, Dec 2016
Jack Clark and Dario Amodei · 2016
Cited alongside, same era.
Deep Reinforcement Learning from Human Preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Trust Region Policy Optimization, April 2017
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2017
Cited alongside, same era.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Specification gaming: the flip side of AI ingenuity, 2020
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Later among the works it cites.
Recursively Summarizing Books with Human Feedback
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Later among the works it cites.
A General Language Assistant as a Laboratory for Alignment, December 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Later among the works it cites.
Learning from Imperfect Demonstrations from Agents with Varying Dynamics, March 2021
Zhangjie Cao and Dorsa Sadigh · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction, November 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
What failure looks like, 2019
Paul Christiano · 2019
Cited alongside, same era.
Kimin Lee, Laura Smith, and Pieter Abbeel
Cited in the paper.
B-Pref: Benchmarking Preference-Based Reinforcement Learning, November 2021b
Kimin Lee, Laura Smith, Anca Dragan, and Pieter Abbeel
Cited in the paper.
A minimal viable product for alignment, March 2022a
Jan Leike
Cited in the paper.
Why I’m excited about AI-assisted human feedback, March 2022b
Jan Leike
Cited in the paper.
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Closest in time.
Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality, January 2022
Songyuan Zhang, Zhangjie Cao, Dorsa Sadigh, and Yanan Sui · 2022
Closest in time.
Imitation Learning by Estimating Expertise of Demonstrators, June 2022
Mark Beliaev, Andy Shih, Stefano Ermon, Dorsa Sadigh, and Ramtin Pedarsani · 2022
Closest in time.
X-Risk Analysis for AI Research, July 2022
Dan Hendrycks and Mantas Mazeika · 2022
Closest in time.