Fetching the paper…
Reading the bibliography…
It is incredibly easy for a system designer to misspecify the objective for an autonomous system ("robot''), thus motivating the desire to have the robot learn the objective from human behavior instead.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Continuous inverse optimal control with locally optimal examples
Sergey Levine and Vladlen Koltun · 2012
Earlier work this paper cites.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Showing versus doing: Teaching by demonstration
Mark K Ho, Michael Littman, James MacGlashan, Fiery Cushman, and Joseph L Austerweil · 2016
Earlier work this paper cites.
Learning robot objectives from physical human interaction
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Faulty reward functions in the wild, Mar 2017
Jack Clark and Dario Amodei · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Cited alongside, same era.
Pragmatic-pedagogic value alignment
Jaime Fisac, Monica A Gates, Jessica B Hamrick, Chang Liu, Dylan Hadfield-Mennell, Malayandi Palaniappan, Dhruv Malik, S Shankar Sastry, Thomas L Griffiths, and Anca D Dragan · 2018
Effectively learning from pedagogical demonstrations
Mark K Ho, Michael L Littman, Fiery Cushman, and Joseph L Austerweil · 2018
Later among the works it cites.
Specification gaming examples in AI, Jun 2018
Victoria Krakovna · 2018
Later among the works it cites.
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Belson, David M Bryson, Nick Cheney, et al · 2018
Later among the works it cites.
An efficient, generalized Bellman update for cooperative inverse reinforcement learning
Dhruv Malik, Malayandi Palaniappan, Jaime Fisac, Dylan Hadfield-Menell, Stuart Russell, and Anca Dragan · 2018
Later among the works it cites.
Optimal cooperative inference
Scott Cheng-Hsin Yang, Yue Yu, Arash Givchi, Pei Wang, Wai Keen Vong, and Patrick Shafto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Legibility and predictability of robot motion
Anca D Dragan, Kenton CT Lee, and Siddhartha S Srinivasa
Cited in the paper.
Teleoperation with intelligent and customizable interfaces
Anca D. Dragan, Siddhartha S. Srinivasa, and Kenton C. T. Lee
Cited in the paper.
Generalizing the theory of cooperative inference
Pei Wang, Pushpi Paranamana, and Patrick Shafto · 2019
Closest in time.