Fetching the paper…
Reading the bibliography…
Herein we describe our approach to artificial intelligence research, which we call the Alberta Plan.
Licklider, J. C. R. (1960). Man-computer symbiosis
1960
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., Precup, D. (2011) · 1999
Earlier work this paper cites.
A survey of algorithmic methods for partially observed Markov decision processes
Lovejoy, W. S. (1991) · 2005
Earlier work this paper cites.
2012
Cited alongside, same era.
Lecture 6.5–RMSProp
Tieleman, T., Hinton, G. (2012) · 2013
Cited alongside, same era.
Kingma, D., Ba, J. (2014). Adam: A method for stochastic optimization. ArXiv:1412.6980
2014
Cited alongside, same era.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988a)
Cited in the paper.
Prioritized sweeping: Reinforcement learning with less data and less real time
Moore, A. W., Atkeson, C. G. (1993) · 2019
Later among the works it cites.
Efficient learning and planning within the Dyna framework
Peng, J., Williams, R. J. (1993) · 2020
Later among the works it cites.
What’s a good prediction? Challenges in evaluating an agent’s knowledge
Kearney, A., Koop, A., Pilarski, P. M. (2022) · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…