Fetching the paper…
Reading the bibliography…
Traditional reinforcement learning (RL) environments typically are the same for both the training and testing phases.
Hyper-parameter sweep on alphazero general
Wang, H., Emmerich, M., Preuss, M., and Plaat, A · 1903
Earlier work this paper cites.
Xxii. programming a computer for playing chess
Shannon, C. E · 1950
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Samuel, A. L · 1959
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G · 1994
Earlier work this paper cites.
Deep blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h · 2002
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C · 2006
Earlier work this paper cites.
Monte-carlo tree search: A new framework for game ai
Chaslot, G., Bakkes, S., Szita, I., and Spronck, P · 2008
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
The impact of determinism on learning atari 2600 games
Hausknecht, M. J. and Stone, P · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Elf opengo: An analysis and open reimplementation of alphazero
Tian, Y., Ma, J., Gong, Q., Sengupta, S., Chen, Z., Pinkerton, J., and Zitnick, L · 2019
Later among the works it cites.
Alternative loss functions in alphazero-like self-play
Wang, H., Emmerich, M., Preuss, M., and Plaat, A · 2019
Later among the works it cites.
Adversarial examples: Attacks and defenses for deep learning
Yuan, X., He, P., Zhu, Q., and Li, X · 2019
Later among the works it cites.
Monte-carlo tree search as regularized policy optimization
Grill, J.-B., Altché, F., Tang, Y., Hubert, T., Valko, M., Antonoglou, I., and Munos, R · 2020
Later among the works it cites.
Minimalistic attacks: How little it takes to fool deep reinforcement learning policies
Qu, X., Sun, Z., Ong, Y. S., Gupta, A., and Wei, P · 2020
Later among the works it cites.
Increasing generality in machine learning through procedural content generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
Adp with mcts algorithm for gomoku
Tang, Z., Zhao, D., Shao, K., and Lv, L · 2016
Cited alongside, same era.
Alphazero–what’s missing?
Bratko, I · 2018
Cited alongside, same era.
A0c: Alpha zero in continuous action space
Moerland, T. M., Broekens, J., Plaat, A., and Jonker, C. M · 2018
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2019
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017a
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D
Cited in the paper.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al
Cited in the paper.
Risi, S. and Togelius, J · 2020
Later among the works it cites.
Train on small, play the large: Scaling up board games with alphazero and gnn
Ben-Assayag, S. and El-Yaniv, R · 2021
Later among the works it cites.
Convex regularization in monte-carlo tree search
Dam, T. Q., D’Eramo, C., Peters, J., and Pajarinen, J · 2021
Later among the works it cites.
Transfer of fully convolutional policy-value networks between games and game variants
Soemers, D. J., Mella, V., Piette, E., Stephenson, M., Browne, C., and Teytaud, O · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents, 2021
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., and Czarnecki, W. M · 2021
Later among the works it cites.