Fetching the paper…
Reading the bibliography…
Learning policies via preference-based reward learning is an increasingly popular method for customizing agent behavior, but has been shown anecdotally to be prone to spurious correlations and reward hacking behaviors.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Variational inference using implicit distributions
Ferenc Huszár · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al · 2017
Earlier work this paper cites.
Batch active preference-based learning of reward functions
Erdem Biyik and Dorsa Sadigh · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim De Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
Assistive gym: A physics simulation framework for assistive robotics
Zackory Erickson, Vamsee Gangaram, Ariel Kapusta, C Karen Liu, and Charles C Kemp · 2020
Cited alongside, same era.
Quantifying differences in reward functions
Adam Gleave, Michael D Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2020
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity
Policy gradient bayesian robust optimization for imitation learning
Zaynah Javed, Daniel S Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca Dragan, and Ken Goldberg · 2021
Later among the works it cites.
The boltzmann policy distribution: Accounting for systematic suboptimality in human models
Cassidy Laidlaw and Anca Dragan · 2021
Later among the works it cites.
B-pref: Benchmarking preference-based reinforcement learning
Kimin Lee, Laura Smith, Anca Dragan, and Pieter Abbeel · 2021
Later among the works it cites.
Goal misgeneralization in deep reinforcement learning
Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, et al · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Cited alongside, same era.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Causal imitation learning with unobserved confounders
Junzhe Zhang, Daniel Kumor, and Elias Bareinboim · 2020
Cited alongside, same era.
Feature expansive reward learning: Rethinking human input
Andreea Bobu, Marius Wiggert, Claire Tomlin, and Anca D Dragan · 2021
Cited alongside, same era.
Jerry Zhi-Yang He and Anca D Dragan · 2021
Cited alongside, same era.
Safe imitation learning via fast bayesian reward inference from preferences
Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum
Cited in the paper.
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Closest in time.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Closest in time.
Causal imitation learning under temporally correlated noise
Gokul Swamy, Sanjiban Choudhury, J Andrew Bagnell, and Zhiwei Steven Wu · 2022
Closest in time.
Dynamics-aware comparison of learned reward functions
Blake Wulfe, Logan Michael Ellis, Jean Mercat, Rowan Thomas McAllister, and Adrien Gaidon · 2022
Closest in time.
Efficient preference-based reinforcement learning using learned dynamics models
Yi Liu, Gaurav Datta, Ellen Novoseller, and Daniel S Brown · 2023
Closest in time.
Benchmarks and algorithms for offline preference-based reward learning
Daniel Shin, Anca Dragan, and Daniel S Brown · 2023
Closest in time.