Fetching the paper…
Reading the bibliography…
We provide a formal definition of depth-limited games together with an accessible and rigorous explanation of the underlying concepts, both of which were previously missing in imperfect-information games.
R. Selten, Reexamination of the perfectness concept for equilibrium points in extensive games, Economics (1974)
1974
Earlier work this paper cites.
M. L. Littman, Markov games as a framework for multi-agent reinforcement learning, in: Machine learning proceedings 1994, Elsevier, 1994, pp. 157–163
1994
Earlier work this paper cites.
H. L. Cole, N. Kocherlakota, Dynamic games with hidden actions and hidden states, Journal of Economic Theory 98 (1) (2001) 114–126
2001
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, Y. Freund, R. E. Schapire, The nonstochastic multiarmed bandit problem, SIAM journal on computing 32 (1) (2002) 48–77
2002
Earlier work this paper cites.
J. Hu, M. P. Wellman, Nash q q -learning for general-sum stochastic games, Journal of machine learning research 4 (Nov) (2003) 1039–1069
2003
Earlier work this paper cites.
A. Greenwald, K. Hall, R. Serrano, et al., Correlated Q-learning, in: ICML, Vol. 3, 2003, pp. 242–249
2003
Earlier work this paper cites.
E. A. Hansen, D. S. Bernstein, S. Zilberstein, Dynamic programming for partially observable stochastic games, in: AAAI, Vol. 4, 2004, pp. 709–715
2004
Earlier work this paper cites.
M. Buro, Solving the oshi-zumo game, in: Advances in Computer Games, Springer, 2004, pp. 361–366
2004
Earlier work this paper cites.
F. Oliehoek, N. Vlassis, et al., Dec-POMDPs and extensive form games: equivalence of models and algorithms, Ias technical report IAS-UVA-06-02, University of Amsterdam, Intelligent Systems Lab, Amsterdam, The Netherlands (2006)
2006
Earlier work this paper cites.
R. Zarick, B. Pellegrino, N. Brown, C. Banister, Unlocking the potential of deep counterfactual value networks (2020) · 2007
Earlier work this paper cites.
M. Zinkevich, M. Johanson, M. Bowling, C. Piccione, Regret minimization in games with incomplete information, in: Advances in neural information processing systems, 2008, pp. 1729–1736
2008
Earlier work this paper cites.
Y. Shoham, K. Leyton-Brown, Multiagent systems: Algorithmic, game-theoretic, and logical foundations, Cambridge University Press, 2008
2008
Earlier work this paper cites.
F. A. Oliehoek, Decentralized POMDPs, in: Reinforcement Learning, Springer, 2012, pp. 471–503
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
N. Burch, M. Lanctot, D. Szafron, R. Gibson, Efficient Monte Carlo counterfactual regret minimization in games with many player actions, Advances in neural information processing systems 25 (2012)
2012
Earlier work this paper cites.
M. Dermed, L. Charles, Value methods for efficiently solving stochastic games of complete and incomplete information, Ph.D. thesis, Georgia Institute of Technology (2013)
2013
Earlier work this paper cites.
F. A. Oliehoek, Sufficient plan-time statistics for decentralized POMDPs, in: Twenty-Third International Joint Conference on Artificial Intelligence, 2013, pp. 302–308
2013
Earlier work this paper cites.
N. Burch, M. Johanson, M. Bowling, Solving imperfect information games using decomposition., in: AAAI, 2014, pp. 602–608
2014
Cited alongside, same era.
V. Lisý, Alternative selection functions for information set Monte Carlo tree search, Acta Polytechnica 54 (5) (2014) 333–340
2014
Cited alongside, same era.
O. Tammelin, CFR+, CoRR, abs/1407.5042 (2014)
2014
Cited alongside, same era.
2014
Cited alongside, same era.
J. Heinrich, M. Lanctot, D. Silver, Fictitious self-play in extensive-form games., in: ICML, 2015, pp. 805–813
2015
Cited alongside, same era.
2019
Closest in time.
N. Brown, T. Sandholm, Superhuman AI for multiplayer poker, Science 365 (6456) (2019) 885–890
2019
Closest in time.
A. Celli, A. Marchesi, T. Bianchi, N. Gatti, Learning to correlate in multi-player general-sum sequential games, Advances in Neural Information Processing Systems 32 (2019) 13076–13086
2019
Closest in time.
M. Šustr, V. Kovařík, V. Lisý, Monte Carlo continual resolving for online strategy computation in imperfect information games, in: Proceedings of the 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2019, pp. 224–232
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. J. Wiggers, F. A. Oliehoek, D. M. Roijers, Structure in the value function of two-player zero-sum games of incomplete information, in: Proceedings of the Twenty-second European Conference on Artificial Intelligence, IOS Press, 2016, pp. 1628–1629
2016
Cited alongside, same era.
S. K. Jakobsen, T. B. Sørensen, V. Conitzer, Timeability of extensive-form games, in: Proceedings of the 2016 ACM Conference on Innovations in Theoretical Computer Science, 2016, pp. 191–199
2016
Cited alongside, same era.
J. Y. Halpern, R. Pass, Sequential equilibrium in games of imperfect recall., in: KR, 2016, pp. 278–287
2016
Cited alongside, same era.
J. S. Dibangoye, C. Amato, O. Buffet, F. Charpillet, Optimally solving Dec-POMDPs as continuous-state MDPs, Journal of Artificial Intelligence Research 55 (2016) 443–497
2016
Cited alongside, same era.
M. Moravčík, M. Schmid, N. Burch, V. Lisý, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, M. Bowling, Deepstack: Expert-level artificial intelligence in heads-up no-limit poker, Science 356 (6337) (2017) 508–513
2017
Cited alongside, same era.
N. Brown, T. Sandholm, Superhuman AI for heads-up no-limit poker: Libratus beats top professionals, Science (2017) eaao1733
2017
Cited alongside, same era.
N. Burch, Time and space: Why imperfect information games are hard, Ph.D. thesis, University of Alberta (2017)
2017
Cited alongside, same era.
2020
Closest in time.
2020
Closest in time.
Gtlib2 implementation of DL-CFR-NN, https://gitlab.fel.cvut.cz/game-theory-aic/GTLib2/-/blob/master/algorithms/cfr_dl.cpp , accessed: 2020-09-17 (2020)
2020
Closest in time.
B. Zhang, T. Sandholm, Small nash equilibrium certificates in very large games, Advances in Neural Information Processing Systems 33 (2020) 7161–7172
2020
Closest in time.
2020
Closest in time.
Loss functions for DL-CFR-NN, https://gitlab.fel.cvut.cz/seitzdom/value_func_repo_python/-/blob/master/nn/loss_functions.py , accessed: 2020-10-20 (2020)
2020
Closest in time.
V. Kovařík, M. Schmid, N. Burch, M. Bowling, V. Lisý, Rethinking formal models of partially observable multiagent decision making, Artificial Intelligence (2021) 103645
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
B. Zhang, T. Sandholm, Finding and certifying (near-) optimal strategies in black-box extensive-form games, in: AAAI Workshop on Reinforcement Learning in Games, 2021, pp. 1–10
2021
Closest in time.
K. Horák, B. Bošanský, Solving partially observable stochastic games with public observations, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, 2019, pp. 2029–2036
2036
Closest in time.