Fetching the paper…
Reading the bibliography…
We consider the task of building strong but human-like policies in multi-agent decision-making problems, given examples of human behavior.
Non-cooperative games
Nash, J · 1951
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
Blackwell, D. et al · 1956
Earlier work this paper cites.
Approximation to bayes risk in repeated play
Hannan, J · 1957
Earlier work this paper cites.
Evolutionary stable strategies and game dynamics
Taylor, P. D. and Jonker, L. B · 1978
Earlier work this paper cites.
Evolution and the Theory of Games
Smith, J. M · 1982
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K · 1994
Earlier work this paper cites.
Quantal response equilibria for normal form games
McKelvey, R. D. and Palfrey, T. R · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E · 1997
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Deep Blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h · 2002
Earlier work this paper cites.
Beating the adaptive bandit with high probability
Abernethy, J. and Rakhlin, A · 2009
Earlier work this paper cites.
Lecture notes on online learning, 2009
Rakhlin, A · 2009
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
Boularias, A., Kober, J., and Peters, J · 2011
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Better computer go player with neural network and long-term prediction, 2016
Tian, Y. and Zhu, Y · 2016
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2017
Earlier work this paper cites.
Residual networks for computer go
Cazenave, T · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M · 2017
Cited alongside, same era.
Beyond monte carlo tree search: Playing go with deep alternative neural network and long-term evaluation
Wang, J., Wang, W., Wang, R., and Gao, W · 2017
Cited alongside, same era.
Emulating human play in a leading mobile card game
Baier, H., Sattaur, A., Powley, E., Devlin, S., Rollason, J., and Cowling, P · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Later among the works it cites.
Learning to play no-press diplomacy with best response policy iteration
Anthony, T., Eccles, T., Tacchetti, A., Kramár, J., Gemp, I., Hudson, T., Porcel, N., Lanctot, M., Perolat, J., Everett, R., Singh, S., Graepel, T., and Bachrach, Y · 2020
Later among the works it cites.
The hanabi challenge: A new frontier for ai research
Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., et al · 2020
Later among the works it cites.
The game is not over yet—go in the post-alphago era
Egri-Nagy, A. and Törmänen, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2018
Cited alongside, same era.
Forum post on alphazero news (post by user matthewlai)
Lai, M · 2018
Cited alongside, same era.
What game are we playing? end-to-end learning in normal and extensive form games
Ling, C. K., Fang, F., and Kolter, J. Z · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Go neural net sandbox
Wu, D · 2018
Cited alongside, same era.
Superhuman AI for multiplayer poker
Brown, N. and Sandholm, T · 2019
Cited alongside, same era.
Gray, J., Lerer, A., Bakhtin, A., and Brown, N · 2020
Later among the works it cites.
Monte-carlo tree search as regularized policy optimization
Grill, J.-B., Altché, F., Tang, Y., Hubert, T., Valko, M., Antonoglou, I., and Munos, R · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
Hu, H., Lerer, A., Peysakhovich, A., and Foerster, J · 2020
Later among the works it cites.
Leela chess zero information page on "neural network topology"
LC0 · 2020
Later among the works it cites.
Improving policies via search in cooperative partially observable games
Lerer, A., Hu, H., Foerster, J., and Brown, N · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Later among the works it cites.
Accelerating self-play learning in go
Wu, D · 2020
Later among the works it cites.
Introducing maia, a human-like neural network chess engine, 2021
Anderson, A., McIlroy-Young, R., Sen, S., and Kleinberg, J · 2021
Closest in time.
No-press diplomacy from scratch
Bakhtin, A., Wu, D., Lerer, A., and Brown, N · 2021
Closest in time.
Fast policy extragradient methods for competitive games with entropy regularization
Cen, S., Wei, Y., and Chi, Y · 2021
Closest in time.
K-level resoning for zero-shot coordination in hanabi
Cui, B., Hu, H., Pineda, L., and Foerster, J · 2021
Closest in time.
Scalable online planning via reinforcement learning fine-tuning
Fickinger, A., Hu, H., Amos, B., Russell, S., and Brown, N · 2021
Closest in time.
On pathologies in kl-regularized reinforcement learning from expert demonstrations
Rudner, T. G., Lu, C., Osborne, M., Gal, Y., and Teh, Y. W · 2021
Closest in time.
Evaluation of human-ai teams for learned and rule-based agents in hanabi
Siu, H. C., Peña, J., Chang, K. C., Chen, E., Zhou, Y., Lopez, V. J., Palko, K., and Allen, R. E · 2021
Closest in time.