Fetching the paper…
Reading the bibliography…
A* is a popular path-finding algorithm, but it can only be applied to those domains where a good heuristic function is known.
A formal basis for the heuristic determination of minimum cost paths
Hart, P. E., Nilsson, N. J., and Raphael, B. (1968) · 1968
Earlier work this paper cites.
The expected-outcome model of two-player games
Abramson, B. D. (1987) · 1987
Earlier work this paper cites.
Fibonacci heaps and their uses in improved network optimization algorithms
Fredman, M. L. and Tarjan, R. E. (1987) · 1987
Earlier work this paper cites.
Monte carlo go
Brügmann, B. (1993) · 1993
Earlier work this paper cites.
Modification of uct with patterns in monte-carlo go
Gelly, S., Wang, Y., Teytaud, O., Patterns, M. U., and Tao, P. (2006) · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015) · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Xu, B., Wang, N., Chen, T., and Li, M. (2015) · 2015
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017a)
Cited in the paper.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017b)
Cited in the paper.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Later among the works it cites.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D. (2017) · 2017
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O. (2018) · 2018
Closest in time.
Observe and look further: Achieving consistent performance on atari
Pohlen, T., Piot, B., Hester, T., Azar, M. G., Horgan, D., Budden, D., Barth-Maron, G., van Hasselt, H., Quan, J., Večerík, M., et al. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Closest in time.