Fetching the paper…
Reading the bibliography…
In batch reinforcement learning, there can be poorly explored state-action pairs resulting in poorly learned, inaccurate models and poorly performing associated policies.
Networks for approximation and learning
Poggio, T. and Girosi, F · 1990
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Li, L., Walsh, T. J., and Littman, M. L · 2006
Earlier work this paper cites.
Adaptive filtering prediction and control
Goodwin, G. C. and Sin, K. S · 2014
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
Jiang, N., Kulesza, A., Singh, S., and Lewis, R · 2015
Cited alongside, same era.
Mitigating planner overfitting in model-based reinforcement learning
Arumugam, D., Abel, D., Asadi, K., Gopalan, N., Grimm, C., Lee, J. K., Lehnert, L., and Littman, M · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Discount factor as a regularizer in reinforcement learning
Amit, R., Meir, R., and Ciosek, K · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…