Fetching the paper…
Reading the bibliography…
To convey desired behavior to a Reinforcement Learning (RL) agent, a designer must choose a reward function for the environment, arguably the most important knob designers have in interacting with RL agents.
Complexity analysis of real-time reinforcement learning
Koenig, S. and Simmons, R. G. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, S. J. and Norvig, P. (1994) · 1994
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. (1998) · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. (1999) · 1999
Earlier work this paper cites.
A near-optimal polynomial time algorithm for learning in certain classes of stochastic games
Brafman, R. I. and Tennenholtz, M. (2000) · 2000
Cited alongside, same era.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. (2000) · 2000
Cited alongside, same era.
Effective reinforcement learning for mobile robots
Smart, W. D. and Kaelbling, L. P. (2002) · 2002
Cited alongside, same era.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Cited alongside, same era.
Apprenticeship learning using linear programming
Syed, U., Bowling, M., and Schapire, R. E. (2008) · 2008
Cited alongside, same era.
Action-gap phenomenon in reinforcement learning
Farahmand, A.-m. (2011) · 2011
Later among the works it cites.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R. (2015) · 2015
Later among the works it cites.
The dependence of effective planning horizon on model accuracy
Jiang, N., Kulesza, A., Singh, S., and Lewis, R. (2015) · 2015
Later among the works it cites.
On value function representation of long horizon problems
Lehnert, L., Laroche, R., and van Seijen, H. (2018) · 2018
Later among the works it cites.
On the expressivity of Markov reward
Abel, D., Dabney, W., Harutyunyan, A., Ho, M. K., Littman, M. L., Precup, D., and Singh, S. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…