Fetching the paper…
Reading the bibliography…
Learning robot objective functions from human input has become increasingly important, but state-of-the-art techniques assume that the human's desired objective lies within the robot's hypothesis space.
Theory of games and economic behavior
J. Von Neumann and O. Morgenstern · 1945
Earlier work this paper cites.
When Is a Linear Control System Optimal?
R. E. Kalman · 1964
Earlier work this paper cites.
Impedance control: An approach to manipulation: Part ii–implementation
N. Hogan · 1985
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Ng and S. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Goal inference as inverse planning
C. L. Baker, J. B. Tenenbaum, and R. R. Saxe · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Cited alongside, same era.
Shared autonomy via hindsight optimization
S. Javdani, S. S. Srinivasa, and J. A. Bagnell · 2015
Cited alongside, same era.
Learning preferences for manipulation tasks from online coactive feedback
A. Jain, S. Sharma, T. Joachims, and A. Saxena · 2015
Cited alongside, same era.
Movement primitives via optimization
A. D. Dragan, K. Muelling, J. A. Bagnell, and S. S. Srinivasa · 2015
Cited alongside, same era.
Learning robot objectives from physical human interaction
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
Inverse reward design
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. D. Dragan · 2017
Later among the works it cites.
S. Milli, D. Hadfield-Menell, A. Dragan, and S. Russell · 2017
Later among the works it cites.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, J. Peters, et al · 2018
Closest in time.
Learning from physical human corrections, one feature at a time
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2018
Closest in time.
Probabilistically safe robot planning with confidence-based human predictions
J. F. Fisac, A. Bajcsy, S. L. Herbert, D. Fridovich-Keil, S. Wang, C. J. Tomlin, and A. D. Dragan · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Variational inverse control with events: A general framework for data-driven reward definition
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine
Cited in the paper.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir
Cited in the paper.
Finding locally optimal, collision-free trajectories with sequential convex optimization
J. Schulman, J. Ho, A. Lee, I. Awwal, H. Bradlow, and P. Abbeel
Cited in the paper.