Fetching the paper…
Reading the bibliography…
Real-world robotic tasks require complex reward functions.
Active mobile robot localization
W. Burgard, D. Fox, and S. Thrun · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Maximum margin planning
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W. B. Knox and P. Stone · 2008
Earlier work this paper cites.
Active learning for reward estimation in inverse reinforcement learning
M. Lopes, F. Melo, and L. Montesano · 2009
Earlier work this paper cites.
Where do rewards come from
S. Singh, R. L. Lewis, and A. G. Barto · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel · 2011
Earlier work this paper cites.
Infinite time horizon maximum causal entropy inverse reinforcement learning
M. Bloem and N. Bambos · 2014
Earlier work this paper cites.
Active reward learning
C. Daniel, M. Viering, J. Metz, O. Kroemer, and J. Peters · 2014
Earlier work this paper cites.
Active advice seeking for inverse reinforcement learning
P. Odom and S. Natarajan · 2015
Earlier work this paper cites.
Learning preferences for manipulation tasks from online coactive feedback
A. Jain, S. Sharma, T. Joachims, and A. Saxena · 2015
Earlier work this paper cites.
Active reward learning with a novel acquisition function
C. Daniel, O. Kroemer, M. Viering, J. Metz, and J. Peters · 2015
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Planning for autonomous cars that leverage effects on human actions
D. Sadigh, S. Sastry, S. A. Seshia, and A. D. Dragan · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, A. Dragan, P. Abbeel, and S. Russell · 2016
Cited alongside, same era.
Inverse reward design
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Cited alongside, same era.
Scalable end-to-end autonomous vehicle testing via rare-event simulation
M. O’Kelly, A. Sinha, H. Namkoong, J. Duchi, and R. Tedrake · 2018
Later among the works it cites.
Rigorous agent evaluation: An adversarial approach to uncover catastrophic failures
J. Uesato, A. Kumar, C. Szepesvari, T. Erez, A. Ruderman, K. Anderson, N. Heess, P. Kohli, et al · 2018
Later among the works it cites.
Simplifying reward design through divide-and-conquer
E. Ratner, D. Hadfield-Menell, and A. D. Dragan · 2018
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Later among the works it cites.
Ray: A distributed framework for emerging { \{ AI } \} applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning robot objectives from physical human interaction
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2017
Cited alongside, same era.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Cited alongside, same era.
Deep bayesian active learning with image data
Y. Gal, R. Islam, and Z. Ghahramani · 2017
Cited alongside, same era.
Carla: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Cited alongside, same era.
Learning to parse natural language to grounded reward functions with weak supervision
E. C. Williams, N. Gopalan, M. Rhee, and S. Tellex · 2018
Cited alongside, same era.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Cited alongside, same era.
From language to goals: Inverse reinforcement learning for vision-based instruction following
J. Fu, A. Korattikara, S. Levine, and S. Guadarrama · 2019
Later among the works it cites.
Risk-aware active inverse reinforcement learning
D. S. Brown, Y. Cui, and S. Niekum · 2019
Later among the works it cites.
Learning human objectives by evaluating hypothetical behavior, 2019
S. Reddy, A. D. Dragan, S. Levine, S. Legg, and J. Leike · 2019
Later among the works it cites.
Tri taking on the hard problems in manipulation research toward making human-assist robots reliable and robust
R. Tedrake · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Later among the works it cites.
The empathic framework for task learning from implicit human feedback
Y. Cui, Q. Zhang, A. Allievi, P. Stone, S. Niekum, and W. B. Knox · 2020
Later among the works it cites.
Preference-based learning for exoskeleton gait optimization
M. Tucker, E. Novoseller, C. Kann, Y. Sui, Y. Yue, J. W. Burdick, and A. D. Ames · 2020
Later among the works it cites.
Generative modeling of environments with scene grammars and variational inference
G. Izatt and R. Tedrake · 2020
Later among the works it cites.