Fetching the paper…
Reading the bibliography…
Designing a good reward function is essential to robot planning and reinforcement learning, but it can also be challenging and frustrating.
Algorithms for Inverse Reinforcement Learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
MCMC for Doubly-intractable Distributions
Iain Murray, Zoubin Ghahramani, and David J. C. MacKay · 2006
Earlier work this paper cites.
A Game-theoretic Approach to Apprenticeship Learning
Umar Syed and Robert E. Schapire · 2007
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Where do rewards come from? , pages 111–116
Satinder Singh, Richard L. Lewis, , and Andrew G. Barto · 2010
Earlier work this paper cites.
Planning Human-Aware Motions using a Sampling-based Costmap Planner
J. Mainprice, E. Akin Sisbot, L. Jaillet, J. Cortés, R. Alami, and T. Siméon · 2011
Earlier work this paper cites.
Preference-learning based Inverse Reinforcement Learning for Dialog Control
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami · 2012
Cited alongside, same era.
A Strategy-Aware Technique for Learning Behaviors from Discrete Human Feedback
Robert Tyler Loftin, James MacGlashan, Bei Peng, Matthew E Taylor, Michael L Littman, Jeff Huang, and David L Roberts · 2014
Cited alongside, same era.
Motion Planning with Sequential Convex Optimization and Convex Collision Checking
John Schulman, Yan Duan, Jonathan Ho, Alex Lee, Ibrahim Awwal, Henry Bradlow, Jia Pan, Sachin Patil, Ken Goldberg, and Pieter Abbeel · 2014
Cited alongside, same era.
Active Reward Learning with a Novel Acquisition Function
Christian Daniel, Oliver Kroemer, Malte Viering, Jan Metz, and Jan Peters · 2015
Cited alongside, same era.
Learning Preferences for Manipulation Tasks from Online Coactive Feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena · 2015
Cited alongside, same era.
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Later among the works it cites.
Learning the Preferences of Ignorant, Inconsistent Agents
Owain Evans, Andreas Stuhlmüller, and Noah D. Goodman · 2016
Later among the works it cites.
Deep Reinforcement Learning from Human Preferences
Paul Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Later among the works it cites.
Inverse Reward Design
Dylan Hadfield-Menell, Smitha Milli, Stuart Russell, Pieter Abbeel, and Anca D. Dragan · 2017
Later among the works it cites.
Active Preference-based Learning of Reward Functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Later among the works it cites.
Design and Evaluation of Adverb Palette: A GUI for Selecting Tradeoffs in Multi-objective Optimization Problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grounding English Commands to Reward Functions
James MacGlashan, Monica Babes-Vroman, Marie desJardins, Michael L. Littman, Smaranda Muresan, Shawn Squire, Stefanie Tellex, Dilip Arumugam, and Lei Yang · 2015
Cited alongside, same era.
Meher T Shaikh and Michael A Goodrich · 2017
Later among the works it cites.