Fetching the paper…
Reading the bibliography…
We attack the state-of-the-art Go-playing AI system KataGo by training adversarial policies against it, achieving a >97% win rate against KataGo running at superhuman settings.
Iterative solution of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
Life in the game of Go
Benson, D. B · 1976
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 1976
Earlier work this paper cites.
Searching for Solutions in Games and Artificial Intelligence
Allis, L. V · 1994
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search
Coulom, R · 2007
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
Johanson, M., Waugh, K., Bowling, M. H., and Zinkevich, M · 2011
Earlier work this paper cites.
Multi-armed bandits with episode context
Rosin, C. D · 2011
Earlier work this paper cites.
PACHI: State of the art open source Go program
Baudiš, P. and Gailly, J.-l · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. J., and Fergus, R · 2014
Earlier work this paper cites.
The game of Go, 2014
Tromp, J · 2014
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Adversarial attacks on neural network policies
Huang, S. H., Papernot, N., Goodfellow, I. J., Duan, Y., and Abbeel, P · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Perolat, J., Silver, D., and Graepel, T · 2017
Earlier work this paper cites.
Equilibrium approximation quality of current no-limit poker bots
Lisý, V. and Bowling, M · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Cited alongside, same era.
CIFAR10 to compare visual recognition performance between deep neural networks and humans
Ho-Phuoc, T · 2018
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Player of games
Schmid, M., Moravcik, M., Burch, N., Kadlec, R., Davidson, J., Waugh, K., Bard, N., Timbers, F., Lanctot, M., Holland, Z., Davoodi, E., Christianson, A., and Bowling, M · 2021
Later among the works it cites.
Adversarial policy training against deep reinforcement learning
Wu, X., Guo, W., Wei, H., and Xing, X · 2021
Later among the works it cites.
Constitutional AI: Harmlessness from AI feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Closest in time.
Imitating opponent to win: Adversarial policy imitation learning in two-player competitive games
Bui, T. V., Mai, T., and Nguyen, T. H · 2022
Closest in time.
Go ratings, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Pérolat, J., Jaderberg, M., and Graepel, T · 2019
Cited alongside, same era.
On evaluating adversarial robustness
Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Cited alongside, same era.
Leela Zero, 2019
Pascutto, G.-C · 2019
Cited alongside, same era.
ELF OpenGo: an analysis and open reimplementation of AlphaZero
Tian, Y., Ma, J., Gong, Q., Sengupta, S., Chen, Z., Pinkerton, J., and Zitnick, L · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Cited alongside, same era.
Accelerating self-play learning in Go
Wu, D. J · 2019
Cited alongside, same era.
Coulom, R · 2022
Closest in time.
Reducing exploitability with population based training
Czempin, P. and Gleave, A · 2022
Closest in time.
European Go database, 2022
EGD · 2022
Closest in time.
European pros, 2022
Federation, E. G · 2022
Closest in time.
summarize_sgfs.py, 2022
Haoda, F. and Wu, D. J · 2022
Closest in time.
Challenges and countermeasures for adversarial attacks on deep reinforcement learning
Ilahi, I., Usama, M., Qadir, J., Janjua, M. U., Al-Fuqaha, A., Hoang, D. T., and Niyato, D · 2022
Closest in time.
Are AlphaZero-like agents robust to adversarial perturbations?
Lan, L.-C., Zhang, H., Wu, T.-R., Tsai, M.-Y., Wu, I.-C., and Hsieh, C.-J · 2022
Closest in time.
Mastering the game of Stratego with model-free multiagent reinforcement learning
Perolat, J., de Vylder, B., Hennes, D., Tarassov, E., Strub, F., de Boer, V., Muller, P., Connor, J. T., Burch, N., Anthony, T., McAleer, S., Elie, R., Cen, S. H., Wang, Z., Gruslys, A., Malysheva, A., Khan, M., Ozair, S., Timbers, F., Pohlen, T., Eccles, T., Rowland, M., Lanctot, M., Lespiau, J.-B., Piot, B., Omidshafiei, S., Lockhart, E., Sifre, L., Beauguerlange, N., Munos, R., Silver, D., Singh, S., Hassabis, D., and Tuyls, K · 2022
Closest in time.
NeuralZ06 bot configuration settings, 2022
Rob · 2022
Closest in time.
Approximate exploitability: Learning a best response in large games
Timbers, F., Bard, N., Lockhart, E., Lanctot, M., Schmid, M., Burch, N., Schrittwieser, J., Hubert, T., and Bowling, M · 2022
Closest in time.
KataGo benchmark, 2022
Yao, D · 2022
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Closest in time.