Fetching the paper…
Reading the bibliography…
Researchers have demonstrated that neural networks are vulnerable to adversarial examples and subtle environment changes, both of which one can view as a form of distribution shift.
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, and et al · 1908
Earlier work this paper cites.
Searching for solutions in games and artificial intelligence
V. L. Allis · 1994
Earlier work this paper cites.
Chinook the world man-machine checkers champion
J. Schaeffer, R. Lake, P. Lu, and M. Bryant · 1996
Earlier work this paper cites.
Deep blue
M. Campbell, A. J. Hoane Jr, and F. Hsu · 2002
Earlier work this paper cites.
Dataset Shift in Machine Learning
J. Quionero-Candela, M. Sugiyama, A. Schwaighofer, and N.D. Lawrence · 2009
Earlier work this paper cites.
Information set Monte Carlo tree search
P.I. Cowling, E.J. Powley, and D. Whitehouse · 2010
Earlier work this paper cites.
Pachi: State of the art open source go program
Petr Baudiš and Jean-loup Gailly · 2011
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
M. Johanson, K. Waugh, M. Bowling, and Martin M. Zinkevich · 2011
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
M. Johanson · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Slumbot NL: Solving large games with counterfactual regret minimization using sampling and distributed processing
E.G. Jackson · 2013
Earlier work this paper cites.
Measuring the size of large no-limit poker games
M. Johanson · 2013
Earlier work this paper cites.
Online Monte Carlo counterfactual regret minimization for search in imperfect information games
V. Lisý, M. Lanctot, and M. Bowling · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, and et al · 2015
Cited alongside, same era.
Solving heads-up limit texas hold’em
O. Tammelin, N. Burch, M. Johanson, and M. Bowling · 2015
Cited alongside, same era.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P.F. Christiano, J. Schulman, and D. Mané · 2016
Cited alongside, same era.
Online Agent Modelling in Human-Scale Problems
N. DC Bard · 2016
Cited alongside, same era.
Eqilibrium approximation quality of current no-limit poker bots
Depth-limited solving for imperfect-information games
N. Brown, T. Sandholm, and B. Amos · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, and Julian Schrittwieser et al · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
S. Srinivasan, M. Lanctot, V. Zambaldi, and et al · 2018
Later among the works it cites.
A generalized method for empirical game theoretic analysis
K. Tuyls, J. Perolat, M. Lanctot, J.Z. Leibo, and T. Graepel · 2018
Later among the works it cites.
Measuring and characterizing generalization in deep reinforcement learning
S. Witty, J.K. Lee, E. Tosch, A. Atrey, M.L. Littman, and D.D. Jensen · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Viliam Lisý and Michael H. Bowling · 2016
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2017
Cited alongside, same era.
Solving for best responses and equilibria in extensive-form games with reinforcement learning methods
A. Greenwald, J. Li, and E. Sodomka · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
M. Moravčík, M. Schmid, N. Burch, and et al · 2017
Cited alongside, same era.
Synthesizing robust adversarial examples
A. Athalye, L. Engstrom, A. Ilyas, and K. Kwok · 2018
Cited alongside, same era.
D. Balduzzi, K. Tuyls, J. Pérolat, and T. Graepel · 2018
Cited alongside, same era.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2018
Later among the works it cites.
α \alpha -rank: Multi-agent evaluation by evolution
S. Omidshafiei, C. Papadimitriou, Piliouras G, K. Tuyls, and et al · 2019
Later among the works it cites.
Adversarial policies: Attacking deep reinforcement learning
A. Gleave, M. Dennis, C. Wild, N. Kant, S. Levine, and S. Russell · 2020
Closest in time.
The advantage regret-matching actor-critic
A. Gruslys, M. Lanctot, R. Munos, F. Timbers, M. Schmid, and et al · 2020
Closest in time.
Vector quantized models for planning
S. Ozair, Y. Li, A. Razavi, I. Antonoglou, A. van den Oord, and O. Vinyals · 2021
Closest in time.
Martin Schmid, Matej Moravcik, Neil Burch, Rudolf Kadlec, Josh Davidson, Kevin Waugh, Nolan Bard, Finbarr Timbers, Marc Lanctot, Zach Holland, et al · 2021
Closest in time.