Fetching the paper…
Reading the bibliography…
Search is an important tool for computing effective policies in single- and multi-agent environments, and has been crucial for achieving superhuman performance in several benchmark fully and partially observable games.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G · 1994
Earlier work this paper cites.
Rollout algorithms for stochastic scheduling problems
Bertsekas, D. P. and Castanon, D. A · 1999
Earlier work this paper cites.
Deep Blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h · 2002
Earlier work this paper cites.
Finding approximate pomdp solutions through belief compression
Roy, N., Gordon, G., and Thrun, S · 2005
Earlier work this paper cites.
Online planning algorithms for pomdps
Ross, S., Pineau, J., Paquet, S., and Chaib-Draa, B · 2008
Earlier work this paper cites.
Monte-carlo planning in large pomdps
Silver, D. and Veness, J · 2010
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. and Stone, P · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Evaluating and modelling hanabi-playing agents
Walton-Rivers, J., Williams, P. R., Bartle, R., Perez-Liebana, D., and Lucas, S. M · 2017
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Solving partially observable stochastic games with public observations
Horák, K. and Bošanskỳ, B · 2019
Later among the works it cites.
Rethinking formal models of partially observable multiagent decision making
Kovařík, V., Schmid, M., Burch, N., Bowling, M., and Lisỳ, V · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2019
Later among the works it cites.
The hanabi challenge: A new frontier for ai research
Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., Dunning, I., Mourad, S., Larochelle, H., Bellemare, M. G., and Bowling, M · 2020
Later among the works it cites.
Simplified action decoder for deep multi-agent reinforcement learning
Hu, H. and Foerster, J. N · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, D · 2018
Cited alongside, same era.
Superhuman AI for multiplayer poker
Brown, N. and Sandholm, T · 2019
Cited alongside, same era.
Diverse agents for ad-hoc cooperation in hanabi
Canaan, R., Togelius, J., Nealen, A., and Menzel, S · 2019
Cited alongside, same era.
Bayesian action decoder for deep multi-agent reinforcement learning
Foerster, J., Song, F., Hughes, E., Burch, N., Dunning, I., Whiteson, S., Botvinick, M., and Bowling, M · 2019
Cited alongside, same era.
Later among the works it cites.
“other-play”for zero-shot coordination
Hu, H., Peysakhovich, A., Lerer, A., and Foerster, J · 2020
Later among the works it cites.
Improving policies via search in cooperative partially observable games
Lerer, A., Hu, H., Foerster, J. N., and Brown, N · 2020
Later among the works it cites.
Joint policy search for multi-agent collaboration with imperfect information
Tian, Y., Gong, Q., and Jiang, T · 2020
Later among the works it cites.
Hu, H., Lerer, A., Cui, B., Pineda, L., Wu, D., Brown, N., and Foerster, J. N · 2021
Closest in time.