Fetching the paper…
Reading the bibliography…
Specifying reward functions for robots that operate in environments without a natural reward signal can be challenging, and incorrectly specified rewards can incentivise degenerate or dangerous behavior.
Prospect Theory: An Analysis of Decision Under Risk
Daniel Kahneman and Amos Tversky · 1979
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Computational rationalization: The inverse equilibrium problem
Kevin Waugh, Brian D Ziebart, and J Andrew Bagnell · 2013
Earlier work this paper cites.
Who is (more) rational?
Syngjoo Choi, Shachar Kariv, Wieland Müller, and Dan Silverman · 2014
Earlier work this paper cites.
Learning the Preferences of Ignorant, Inconsistent Agents
Owain Evans, Andreas Stuhlmueller, and Noah D. Goodman · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Learning robot objectives from physical human interaction
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Cited alongside, same era.
Risk-sensitive inverse reinforcement learning via coherent risk models
Anirudha Majumdar, Sumeet Singh, Ajay Mandlekar, and Marco Pavone · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
Model mis-specification and inverse reinforcement learning, Feb 2017
Jacob Steinhardt and Owain Evans · 2017
Sören Mindermann, Rohin Shah, Adam Gleave, and Dylan Hadfield-Menell · 2018
Later among the works it cites.
Where do you think you’re going?: Inferring beliefs about dynamics from behavior
Sid Reddy, Anca Dragan, and Sergey Levine · 2018
Later among the works it cites.
Using natural language for reward shaping in reinforcement learning
Prasoon Goyal, Scott Niekum, and Raymond J Mooney · 2019
Later among the works it cites.
Literal or pedagogic human? analyzing human model misspecification in objective learning
Smitha Milli and Anca D Dragan · 2019
Later among the works it cites.
Preferences implicit in the state of the world
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Specification gaming examples in ai, April 2018
Victoria Krakovna · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
Later among the works it cites.
Quantifying hypothesis space misspecification in learning from human–robot demonstrations and physical corrections
Andreea Bobu, Andrea Bajcsy, Jaime F Fisac, Sampada Deglurkar, and Anca D Dragan · 2020
Later among the works it cites.
Less is more: Rethinking probabilistic models of human behavior
Andreea Bobu, Dexter RR Scobee, Jaime F Fisac, S Shankar Sastry, and Anca D Dragan · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca D Dragan · 2020
Later among the works it cites.