Fetching the paper…
Reading the bibliography…
Planning problems are among the most important and well-studied problems in artificial intelligence.
Some studies in machine learning using the game of checkers
Samuel, A · 1959
Earlier work this paper cites.
An analysis of alpha-beta pruning
Knuth, D. E. and Moore, R. W · 1975
Earlier work this paper cites.
Connectionist learning of expert preferences by comparison training
Tesauro, G · 1988
Earlier work this paper cites.
On optimal game-tree search using rational meta-reasoning
Russell, S. and Wefald, E · 1989
Earlier work this paper cites.
TD-gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G · 1994
Earlier work this paper cites.
Rationality and intelligence
Russell, S · 1995
Earlier work this paper cites.
Knightcap: A chess program that learns by combining td ( λ \lambda ) with game-tree search
Baxter, J., Tridgell, A., and Weaver, L · 1998
Earlier work this paper cites.
The games computers (and people) play
Schaeffer, J · 2000
Earlier work this paper cites.
Temporal difference learning applied to a high-performance game-playing program
Schaeffer, J., Hlynka, M., and Jussila, V · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Cited alongside, same era.
The handbook of brain theory and neural networks
Arbib, M. A · 2003
Cited alongside, same era.
Using abstraction for planning in sokoban
Botea, A., Müller, M., and Schaeffer, J · 2003
Cited alongside, same era.
RSPSA: enhanced parameter optimization in games
Kocsis, L., Szepesvári, C., and Winands, M. H · 2005
Cited alongside, same era.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Cited alongside, same era.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C · 2006
Cited alongside, same era.
Multi-armed bandits with episode context
Rosin, C. D · 2011
Later among the works it cites.
Learning to search better than your teacher
Chang, K.-W., Krishnamurthy, A., Agarwal, A., Daume, H., and Langford, J · 2015
Later among the works it cites.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P · 2015
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D · 2017
Later among the works it cites.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jünger, M., Liebling, T. M., Naddef, D., Nemhauser, G. L., Pulleyblank, W. R., Reinelt, G., Rinaldi, G., and Wolsey, L. A · 2009
Cited alongside, same era.
Bootstrapping from game tree search
Veness, J., Silver, D., Blair, A., and Uther, W · 2009
Cited alongside, same era.
Metareasoning for monte carlo tree search
Hay, N. and Russell, S. J · 2011
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., et al
Cited in the paper.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., et al
Cited in the paper.
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S · 2017
Later among the works it cites.
Learning model-based planning from scratch
Pascanu, R., Li, Y., Vinyals, O., Heess, N., Buesing, L., Racanière, S., Reichert, D., Weber, T., Wierstra, D., and Battaglia, P · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Later among the works it cites.