Fetching the paper…
Reading the bibliography…
It is often very challenging to manually design reward functions for complex, real-world tasks.
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D., and Hester, T · 1904
Earlier work this paper cites.
Theory of Games and Economic Behavior
von Neumann, J. and Morgenstern, O · 1947
Earlier work this paper cites.
Utility Theory for Decision Making
Fishburn, P. C · 1970
Earlier work this paper cites.
Structural estimation of Markov decision processes
Rust, J · 1994
Earlier work this paper cites.
Identification Problems in the Social Sciences
Manski, C. F · 1995
Earlier work this paper cites.
Combinatorial Theory
Aigner, M · 1996
Earlier work this paper cites.
Learning agents for uncertain environments (extended abstract)
Russell, S · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S · 2000
Earlier work this paper cites.
Partial Identification of Probability Distributions
Manski, C. F · 2003
Earlier work this paper cites.
Principled methods for advising reinforcement learning agents
Wiewiora, E., Cottrell, G. W., and Elkan, C · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Nonparametric identification of behavioral responses to counterfactual policy interventions in dynamic discrete decision processes
Aguirregabiria, V · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Ramachandran, D. and Amir, E · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Dynamic discrete choice structural models: A survey
Aguirregabiria, V. and Mira, P · 2009
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, S. and Norvig, P · 2009
Earlier work this paper cites.
Autonomous helicopter aerobatics through apprenticeship learning
Abbeel, P., Coates, A., and Ng, A. Y · 2010
Cited alongside, same era.
Inverse optimal control with linearly-solvable MDPs
Dvijotham, K. and Todorov, E · 2010
Cited alongside, same era.
Partial identification in econometrics
Tamer, E · 2010
Cited alongside, same era.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D · 2010
Cited alongside, same era.
Modeling interaction via the principle of maximum causal entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K · 2010
Cited alongside, same era.
Preference elicitation and inverse reinforcement learning
Rothkopf, C. A. and Dimitrakakis, C · 2011
Cited alongside, same era.
Scalable agent alignment via reward modeling: A research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Reward learning from narrated demonstrations
Tung, H.-Y., Harley, A. W., Huang, L.-K., and Fragkiadaki, K · 2018
Later among the works it cites.
Combining reward information from multiple sources
Krasheninnikov, D., Shah, R., and van Hoof, H · 2019
Later among the works it cites.
The identification zoo: Meanings of identification in econometrics
Lewbel, A · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
APRIL: Active preference learning-based reinforcement learning
Akrour, R., Schoenauer, M., and Sebag, M · 2012
Cited alongside, same era.
Identification in discrete Markov decision models
Srisuma, S · 2015
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P. F., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Repeated inverse reinforcement learning
Amin, K., Jiang, N., and Singh, S. P · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Palan, M., Landolfi, N. C., Shevchuk, G., and Sadigh, D · 2019
Later among the works it cites.
On the feasibility of learning, rather than assuming, human biases for reward inference
Shah, R., Gundotra, N., Abbeel, P., and Dragan, A · 2019
Later among the works it cites.
End-to-end robotic reinforcement learning without reward engineering
Singh, A., Yang, L., Hartikainen, K., Finn, C., and Levine, S · 2019
Later among the works it cites.
From inverse optimal control to inverse reinforcement learning: A historical review
Ab Azar, N., Shahmansoorian, A., and Davoudi, M · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Jeon, H. J., Milli, S., and Dragan, A · 2020
Later among the works it cites.
Iterative interactive reward learning
Koppol, P., Admoni, H., and Simmons, R · 2020
Later among the works it cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Later among the works it cites.
Identifiability in inverse reinforcement learning
Cao, H., Cohen, S. N., and Szpruch, L · 2021
Later among the works it cites.
Reward identification in inverse reinforcement learning
Kim, K., Garg, S., Shiragur, K., and Ermon, S · 2021
Later among the works it cites.
What matters for adversarial imitation learning?
Orsini, M., Raichuk, A., Hussenot, L., Vincent, D., Dadashi, R., Girgin, S., Geist, M., Bachem, O., Pietquin, O., and Andrychowicz, M · 2021
Later among the works it cites.
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
Bıyık, E., Losey, D. P., Palan, M., Landolfi, N. C., Shevchuk, G., and Sadigh, D · 2022
Closest in time.