Fetching the paper…
Reading the bibliography…
Value alignment is essential for building AI systems that can safely and reliably interact with people.
A behavioral model of rational choice
Herbert A. Simon · 1955
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Amos Tversky and Daniel Kahneman · 1974
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Action understanding as inverse planning
Chris L Baker, Rebecca Saxe, and Joshua B Tenenbaum · 2009
Earlier work this paper cites.
Behavioral cloning
Paul Munro, Hannu Toivonen, Geoffrey I. Webb, Wray Buntine, Peter Orbanz, Yee Whye Teh, Pascal Poupart, Claude Sammut, Caude Sammut, Hendrik Blockeel, Dev Rajnarayan, David Wolpert, Wulfram Gerstner, C. David Page, Sriraam Natarajan, and Geoffrey Hinton · 2011
Earlier work this paper cites.
Learning the preferences of ignorant, inconsistent agents
Owain Evans, Andreas Stuhlmüller, and Noah Goodman · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Cited alongside, same era.
Showing versus doing: Teaching by demonstration
Mark K Ho, Michael Littman, James MacGlashan, Fiery Cushman, and Joseph L Austerweil · 2016
Cited alongside, same era.
Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
Robert Loftin, Bei Peng, James MacGlashan, Michael L Littman, Matthew E Taylor, Jeff Huang, and David L Roberts · 2016
Cited alongside, same era.
When humans aren’t optimal: Robots that collaborate with risk-aware humans
Minae Kwon, Erdem Biyik, Aditi Talati, Karan Bhasin, Dylan P Losey, and Dorsa Sadigh · 2020
Cited alongside, same era.
Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources
Falk Lieder and Thomas L Griffiths · 2020
Cited alongside, same era.
Human irrationality: both bad and good for reward inference
Lawrence Chan, Andrew Critch, and Anca Dragan · 2021
Later among the works it cites.
Robust inverse reinforcement learning under transition dynamics mismatch
Luca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Adrian Weller, and Volkan Cevher · 2021
Later among the works it cites.
Cognitive science as a source of forward and inverse models of human decisions for robotics and control
Mark K. Ho and Thomas L. Griffiths · 2022
Later among the works it cites.
People construct simplified mental representations to plan
Mark K Ho, David Abel, Carlos G Correa, Michael L Littman, Jonathan D Cohen, and Thomas L Griffiths · 2022
Later among the works it cites.
The boltzmann policy distribution: Accounting for systematic suboptimality in human models
Cassidy Laidlaw and Anca Dragan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online bayesian goal inference for boundedly rational planning agents
Tan Zhi-Xuan, Jordyn Mann, Tom Silver, Josh Tenenbaum, and Vikash Mansinghka · 2020
Cited alongside, same era.
Modeling the mistakes of boundedly rational agents within a bayesian theory of mind
Arwa Alanqary, Gloria Z Lin, Joie Le, Tan Zhi-Xuan, Vikash K Mansinghka, and Joshua B Tenenbaum · 2021
Cited alongside, same era.
Rational simplification and rigidity in human planning, Mar 2023
Mark K Ho, Jonathan D Cohen, and Tom Griffiths · 2023
Closest in time.