Fetching the paper…
Reading the bibliography…
Lookahead search has been a critical component of recent AI successes, such as in the games of chess, go, and poker.
The optimal control of partially observable markov decision processes
E. J. Sondik · 1971
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
G. Tesauro · 1994
Earlier work this paper cites.
Deep Blue
M. Campbell, A. J. Hoane Jr, and F.-h. Hsu · 2002
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
F. A. Oliehoek, M. T. Spaan, and N. Vlassis · 2008
Earlier work this paper cites.
Information set monte carlo tree search
P. I. Cowling, E. J. Powley, and D. Whitehouse · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
R. Munos · 2014
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
M. Oquab, L. Bottou, I. Laptev, and J. Sivic · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas · 2016
Earlier work this paper cites.
Monte carlo tree search in continuous action spaces with execution uncertainty
T. Yee, V. Lisỳ, and M. H. Bowling · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
M. Moravčík, M. Schmid, N. Burch, V. Lisỳ, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, and M. Bowling · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Simplified action decoder for deep multi-agent reinforcement learning
H. Hu and J. N. Foerster · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2019
Later among the works it cites.
The differentiable cross-entropy method
B. Amos and D. Yarats · 2020
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, Z. D. Guo, and C. Blundell · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multiple-step greedy policies in approximate and online reinforcement learning
Y. Efroni, G. Dalal, B. Scherrer, and S. Mannor · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
A survey on deep transfer learning
C. Tan, F. Sun, T. Kong, W. Zhang, C. Yang, and C. Liu · 2018
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
N. Bard, J. N. Foerster, S. Chandar, N. Burch, M. Lanctot, H. F. Song, E. Parisotto, V. Dumoulin, S. Moitra, E. Hughes, et al · 2020
Later among the works it cites.
Combining deep reinforcement learning and search for imperfect-information games
N. Brown, A. Bakhtin, A. Lerer, and Q. Gong · 2020
Later among the works it cites.
Monte-carlo tree search as regularized policy optimization
J.-B. Grill, F. Altché, Y. Tang, T. Hubert, M. Valko, I. Antonoglou, and R. Munos · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster · 2020
Later among the works it cites.
Improving policies via search in cooperative partially observable games
A. Lerer, H. Hu, J. N. Foerster, and N. Brown · 2020
Later among the works it cites.
Iterative amortized policy optimization
J. Marino, A. Piché, A. D. Ialongo, and Y. Yue · 2020
Later among the works it cites.
Local search for policy iteration in continuous control
J. T. Springenberg, N. Heess, D. Mankowitz, J. Merel, A. Byravan, A. Abdolmaleki, J. Kay, J. Degrave, J. Schrittwieser, Y. Tassa, et al · 2020
Later among the works it cites.
Learned belief search: Efficiently improving policies in partially observable settings, 2021
H. Hu, A. Lerer, N. Brown, and J. N. Foerster · 2021
Closest in time.
Learning and planning in complex action spaces
T. Hubert, J. Schrittwieser, I. Antonoglou, M. Barekatain, S. Schmitt, and D. Silver · 2021
Closest in time.
Online and offline reinforcement learning by planning with a learned model
J. Schrittwieser, T. Hubert, A. Mandhane, M. Barekatain, I. Antonoglou, and D. Silver · 2021
Closest in time.
The bitter lesson
R. Sutton · 2021
Closest in time.