Fetching the paper…
Reading the bibliography…
The process of revising (or constructing) a policy at execution time -- known as decision-time planning -- has been key to achieving superhuman performance in perfect-information games like chess and Go.
Iterative solution of games by fictitious play
G. W. Brown · 1951
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. Nemirovsky and D. Yudin · 1983
Earlier work this paper cites.
Algorithms for Sequential Decision Making
M. L. Littman · 1996
Earlier work this paper cites.
On-line policy improvement using monte-carlo search
G. Tesauro and G. Galperin · 1996
Earlier work this paper cites.
Quantal response equilibria for extensive form games
R. D. McKelvey and T. R. Palfrey · 1998
Earlier work this paper cites.
Deep Blue
M. Campbell, A. Hoane, and F. hsiung Hsu · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
E. A. Hansen, D. S. Bernstein, and S. Zilberstein · 2004
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2007
Earlier work this paper cites.
A tutorial on particle filtering and smoothing: Fifteen years later
A. Doucet and A. Johansen · 2009
Earlier work this paper cites.
A survey of monte carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Optimally solving Dec-POMDPs as continuous-state MDPs
J. S. Dibangoye, C. Amato, O. Buffet, and F. Charpillet · 2013
Earlier work this paper cites.
Decentralized stochastic control with partial history sharing: A common information approach
A. Nayyar, A. Mahajan, and D. Teneketzis · 2013
Earlier work this paper cites.
Sufficient plan-time statistics for decentralized pomdps
F. A. Oliehoek · 2013
Earlier work this paper cites.
Solving imperfect information games using decomposition
N. Burch, M. Johanson, and M. Bowling · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
J. Heinrich and D. Silver · 2016
Cited alongside, same era.
Timeability of extensive-form games
S. K. Jakobsen, T. B. Sørensen, and V. Conitzer · 2016
Cited alongside, same era.
Refining subgames in large imperfect information games
M. Moravcik, M. Schmid, K. Ha, M. Hladik, and S. J. Gaukrodger · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Cited alongside, same era.
Safe and nested subgame solving for imperfect-information games
The hanabi challenge: A new frontier for ai research
N. Bard, J. N. Foerster, S. Chandar, N. Burch, M. Lanctot, H. F. Song, E. Parisotto, V. Dumoulin, S. Moitra, E. Hughes, I. Dunning, S. Mourad, H. Larochelle, M. G. Bellemare, and M. Bowling · 2020
Later among the works it cites.
Combining deep reinforcement learning and search for imperfect-information games
N. Brown, A. Bakhtin, A. Lerer, and Q. Gong · 2020
Later among the works it cites.
Improving policies via search in cooperative partially observable games
A. Lerer, H. Hu, J. Foerster, and N. Brown · 2020
Later among the works it cites.
Joint policy search for multi-agent collaboration with imperfect information
Y. Tian, Q. Gong, and T. Jiang · 2020
Later among the works it cites.
Expert iteration
T. W. Anthony · 2021
Later among the works it cites.
Scalable online planning via reinforcement learning fine-tuning
A. Fickinger, H. Hu, B. Amos, S. Russell, and N. Brown · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Brown and T. Sandholm · 2017
Cited alongside, same era.
DeepStack: Expert-level artificial intelligence in heads-up no-limit poker
M. Moravčík, M. Schmid, N. Burch, V. Lisý, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, and M. Bowling · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Cited alongside, same era.
RLlib: Abstractions for distributed reinforcement learning
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Cited alongside, same era.
Policy gradient search: Online planning and expert iteration without search trees
T. Anthony, R. Nishihara, P. Moritz, T. Salimans, and J. Schulman · 2019
Cited alongside, same era.
Later among the works it cites.
On the role of planning in model-based deep reinforcement learning
J. B. Hamrick, A. L. Friesen, F. Behbahani, A. Guez, F. Viola, S. Witherspoon, T. Anthony, L. H. Buesing, P. Veličković, and T. Weber · 2021
Later among the works it cites.
Learned belief search: Efficiently improving policies in partially observable settings
H. Hu, A. Lerer, N. Brown, and J. N. Foerster · 2021
Later among the works it cites.
From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization
J. Perolat, R. Munos, J.-B. Lespiau, S. Omidshafiei, M. Rowland, P. Ortega, N. Burch, T. Anthony, D. Balduzzi, B. De Vylder, G. Piliouras, M. Lanctot, and K. Tuyls · 2021
Later among the works it cites.
Subgame solving without common knowledge
B. Zhang and T. Sandholm · 2021
Later among the works it cites.
Planning in stochastic environments with a learned model
I. Antonoglou, J. Schrittwieser, S. Ozair, T. K. Hubert, and D. Silver · 2022
Later among the works it cites.
Modeling strong and human-like gameplay with KL-regularized search
A. P. Jacob, D. J. Wu, G. Farina, A. Lerer, H. Hu, A. Bakhtin, J. Andreas, and N. Brown · 2022
Later among the works it cites.
Rethinking formal models of partially observable multiagent decision making
V. Kovarík, M. Schmid, N. Burch, M. Bowling, and V. Lisý · 2022
Later among the works it cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Meta Fundamental AI Research Diplomacy Team (FAIR), A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, A. P. Jacob, M. Komeili, K. Konath, M. Kwon, A. Lerer, M. Lewis, A. H. Miller, S. Mitts, A. Renduchintala, S. Roller, D. Rowe, W. Shi, J. Spisak, A. Wei, D. Wu, H. Zhang, and M. Zijlstra · 2022
Later among the works it cites.
A fine-tuning approach to belief state modeling
S. Sokota, H. Hu, D. J. Wu, J. Z. Kolter, J. N. Foerster, and N. Brown · 2022
Later among the works it cites.
Approximate exploitability: Learning a best response
F. Timbers, N. Bard, E. Lockhart, M. Lanctot, M. Schmid, N. Burch, J. Schrittwieser, T. Hubert, and M. Bowling · 2022
Later among the works it cites.
Mastering the game of No-Press Diplomacy via human-regularized reinforcement learning and planning
A. Bakhtin, D. J. Wu, A. Lerer, J. Gray, A. P. Jacob, G. Farina, A. H. Miller, and N. Brown · 2023
Closest in time.
Opponent-limited online search for imperfect information games
W. Liu, H. Fu, Q. Fu, and Y. Wei · 2023
Closest in time.