Fetching the paper…
Reading the bibliography…
Recent progress in Game AI has demonstrated that given enough data from human gameplay, or experience gained via simulations, machines can rival or surpass the most skilled human players in classic games such as Go, or commercial computer games such as Starcraft.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
Tracking the red queen: Measurements of adaptive progress in co-evolutionary simulations
D. Cliff and G. F. Miller · 1995
Earlier work this paper cites.
Efficient reinforcement learning through symbiotic evolution
D. E. Moriarty and R. Mikkulainen · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
New Methods for Competitive Coevolution
C. D. Rosin and R. K. Belew · 1997
Earlier work this paper cites.
Search in games with incomplete information: A case study using bridge card play
I. Frank and D. Basin · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
A fast elitist non-dominated sorting genetic algorithm for multi-objective optimization: NSGA-II
K. Deb, S. Agrawal, A. Pratap, and T. Meyarivan · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
Completely Derandomized Self-Adaptation in Evolution Strategies
N. Hansen and A. Ostermeier · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
M. Bowling and M. Veloso · 2002
Earlier work this paper cites.
A heuristic approach for solving decentralized-pomdp: Assessment on the pursuit problem
I. Chades, B. Scherrer, and F. Charpillet · 2002
Earlier work this paper cites.
Multi-attribute Decision Making in a Complex Multiagent Environment Using Reinforcement Learning with Selective Perception
S. Paquet, N. Bernier, and B. Chaib-draa · 2004
Earlier work this paper cites.
A Framework for Sequential Planning in Multi-Agent Settings
P. J. Gmytrasiewicz and P. Doshi · 2005
Earlier work this paper cites.
Learning deterministic finite automata with a smart state labeling evolutionary algorithm
S. M. Lucas and T. J. Reynolds · 2005
Earlier work this paper cites.
The epistemology of war gaming
R. C. Rubel · 2006
Earlier work this paper cites.
Evolving robust and specialized car racing skills
J. Togelius and S. M. Lucas · 2006
Earlier work this paper cites.
Generating diverse opponents with multiobjective evolution
A. Agapitos, J. Togelius, S. M. Lucas, J. Schmidhuber, and A. Konstantinidis · 2008
Earlier work this paper cites.
Investigating learning rates for evolution and temporal difference learning
S. M. Lucas · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
A tutorial on particle filtering and smoothing: Fifteen years later
A. Doucet and A. M. Johansen · 2009
Earlier work this paper cites.
Improving Hearthstone AI by Combining MCTS and Supervised Learning Algorithms
M. Świechowski, T. Tajmajer, and A. Janusz · 2009
Earlier work this paper cites.
Agent subset adversarial search for complex non-cooperative domains
V. Lisy, B. Bosansky, R. Vaculin, and M. Pechoucek · 2010
Earlier work this paper cites.
Monte-Carlo planning in large POMDPs
D. Silver and J. Veness · 2010
Earlier work this paper cites.
Computing approximate nash equilibria and robust best-responses using sampling
M. Ponsen, S. De Jong, and M. Lanctot · 2011
Earlier work this paper cites.
A Bayesian approach for learning and planning in partially observable Markov decision processes
S. Ross, J. Pineau, B. Chaib-draa, and P. Kreitmann · 2011
Earlier work this paper cites.
Online planning for multi-agent systems with bounded communication
F. Wu, S. Zilberstein, and X. Chen · 2011
Earlier work this paper cites.
Multiobjective evolutionary algorithms: A survey of the state of the art
A. Zhou, B.-Y. Qu, H. Li, S.-Z. Zhao, P. N. Suganthan, and Q. Zhang · 2011
Earlier work this paper cites.
A Survey of Monte Carlo Tree Search Methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Information Set Monte Carlo Tree Search
P. I. Cowling, E. J. Powley, and D. Whitehouse · 2012
Earlier work this paper cites.
Characteristics of games
G. S. Elias, R. Garfield, and K. R. Gutschera · 2012
Earlier work this paper cites.
QueryPOMDP: POMDP-Based Communication in Multiagent Systems
F. S. Melo, M. T. J. Spaan, and S. J. Witwicki · 2012
Earlier work this paper cites.
Decentralized POMDPs
F. A. Oliehoek · 2012
Earlier work this paper cites.
Playing at the world: A history of simulating wars, people and fantastic adventures, from chess to role-playing games
J. Peterson · 2012
Earlier work this paper cites.
Portfolio greedy search and simulation for large-scale combat in starcraft
D. Churchill and M. Buro · 2013
Earlier work this paper cites.
Recursive Monte Carlo search for imperfect information games
T. Furtak and M. Buro · 2013
Earlier work this paper cites.
Towards imitation of human driving style in car racing games
J. Muñoz, G. Gutierrez, and A. Sanchis · 2013
Earlier work this paper cites.
Rolling horizon evolution versus tree search for navigation in single-player real-time games
D. Perez, S. Samothrakis, S. Lucas, and P. Rohlfshagen · 2013
Earlier work this paper cites.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
J. Schmidhuber · 2013
Earlier work this paper cites.
Search in Imperfect Information Games using Online Monte Carlo Counterfactual Regret Minimization
M. Lanctot, V. Lisy, and M. Bowling · 2014
Earlier work this paper cites.
Learning a Super Mario controller from examples of human play
G. Lee, M. Luo, F. Zambetta, and X. Li · 2014
Earlier work this paper cites.
Dependency-based word embeddings
O. Levy and Y. Goldberg · 2014
Earlier work this paper cites.
Hierarchical adversarial search applied to real-time strategy games
M. Stanescu, N. A. Barriga, and M. Buro · 2014
Earlier work this paper cites.
Natural evolution strategies
D. Wierstra, T. Schaul, T. Glasmachers, Y. Sun, J. Peters, and J. Schmidhuber · 2014
Earlier work this paper cites.
Scalable Planning and Learning for Multiagent POMDPs
C. Amato and F. A. Oliehoek · 2015
Earlier work this paper cites.
ICBS: improved conflict-based search algorithm for multi-agent pathfinding
E. Boyarski, A. Felner, R. Stern, G. Sharon, D. Tolpin, O. Betzalel, and E. Shimony · 2015
Earlier work this paper cites.
Hierarchical portfolio search: Prismata’s robust AI architecture for games with large search spaces
D. Churchill and M. Buro · 2015
Earlier work this paper cites.
Smooth UCT Search in Computer Poker
J. Heinrich and D. Silver · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, and et al · 2015
Cited alongside, same era.
Illuminating search spaces by mapping elites
J.-B. Mouret and J. Clune · 2015
Cited alongside, same era.
Adversarial hierarchical-task network planning for complex real-time games
S. Ontanón and M. Buro · 2015
Cited alongside, same era.
Multimodal optimization by means of evolutionary algorithms
M. Preuss · 2015
Cited alongside, same era.
Conflict-based search for optimal multi-agent pathfinding
G. Sharon, R. Stern, A. Felner, and N. R. Sturtevant · 2015
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz · 2017
Later among the works it cites.
Autonomous agents modelling other agents: A comprehensive survey and open problems
S. V. Albrecht and P. Stone · 2018
Later among the works it cites.
Deceptive games
D. Anderson, M. Stephenson, J. Togelius, C. Salge, J. Levine, and J. Renz · 2018
Later among the works it cites.
Re-evaluating evaluation
D. Balduzzi, K. Tuyls, J. Perolat, and T. Graepel · 2018
Later among the works it cites.
Dopamine: A research framework for deep reinforcement learning
P. S. Castro, S. Moitra, C. Gelada, S. Kumar, and M. G. Bellemare · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
Xgboost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Cited alongside, same era.
Deep Reinforcement Learning for Accelerating the Convergence Rate
J. Fu, Z. Lin, D. Chen, R. Ng, M. Liu, N. Leonard, J. Feng, and T.-S. Chua · 2016
Cited alongside, same era.
Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
J. Heinrich and D. Silver · 2016
Cited alongside, same era.
MCTS/EA hybrid GVGAI players and game difficulty estimation
H. Horn, V. Volz, D. Perez-Liebana, and M. Preuss · 2016
Cited alongside, same era.
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
J. N. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2018
Later among the works it cites.
Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning
J. N. Foerster, F. Song, E. Hughes, N. Burch, I. Dunning, S. Whiteson, M. Botvinick, and M. Bowling · 2018
Later among the works it cites.
Addressing Function Approximation Error in Actor-Critic Methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Later among the works it cites.
Horizon: Facebook’s open source applied reinforcement learning platform
J. Gauci, E. Conti, Y. Liang, K. Virochsiri, Y. He, Z. Kaden, V. Narayanan, X. Ye, Z. Chen, and S. Fujimoto · 2018
Later among the works it cites.
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Playing Multiaction Adversarial Games: Online Evolutionary Planning Versus Tree Search
N. Justesen, T. Mahlmann, S. Risi, and J. Togelius · 2018
Later among the works it cites.
Automated Curriculum Learning by Rewarding Temporally Rare Events
N. Justesen and S. Risi · 2018
Later among the works it cites.
Illuminating generalization in deep reinforcement learning through procedural level generation
N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2018
Later among the works it cites.
Action abstractions for combinatorial multi-armed bandit tree search
R. O. Moraes, J. R. Marino, L. H. Lelis, and M. A. Nascimento · 2018
Later among the works it cites.
The first MicroRTS artificial intelligence competition
S. Ontañón, N. A. Barriga, C. R. Silva, R. O. Moraes, and L. H. Lelis · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
I. Osband, J. Aslanides, and A. Cassirer · 2018
Later among the works it cites.
Recurrent relational networks
R. Palm, U. Paquet, and O. Winther · 2018
Later among the works it cites.
Generating Novice Heuristics for Post-Flop Poker
F. d. M. Silva, J. Togelius, F. Lantz, and A. Nealen · 2018
Later among the works it cites.
How to Win Kaggle Competitions | Data Science and Machine Learning, 2018
Z. U.-H. Usmani · 2018
Later among the works it cites.
Artificial intelligence and games
G. N. Yannakakis and J. Togelius · 2018
Later among the works it cites.
Automatic bridge bidding using deep reinforcement learning
C.-K. Yeh, C.-Y. Hsieh, and H.-T. Lin · 2018
Later among the works it cites.
A Study on Overfitting in Deep Reinforcement Learning
C. Zhang, O. Vinyals, R. Munos, and S. Bengio · 2018
Later among the works it cites.
Graph neural networks: A review of methods and applications
J. Zhou, G. Cui, Z. Zhang, C. Yang, Z. Liu, L. Wang, C. Li, and M. Sun · 2018
Later among the works it cites.
Feudal multi-agent hierarchies for cooperative reinforcement learning
S. Ahilan and P. Dayan · 2019
Later among the works it cites.
Dota 2 with Large Scale Deep Reinforcement Learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Later among the works it cites.
Superhuman AI for multiplayer poker
N. Brown and T. Sandholm · 2019
Later among the works it cites.
Robust Continuous Build-Order Optimization in StarCraft
D. Churchill, M. Buro, and R. Kelly · 2019
Later among the works it cites.
Learning Local Forward Models on Unforgiving Games
A. Dockhorn, S. M. Lucas, V. Volz, I. Bravi, R. D. Gaina, and D. Perez-Liebana · 2019
Later among the works it cites.
Subgoal-based temporal abstraction in Monte-Carlo tree search
T. Gabor, J. Peter, T. Phan, C. Meyer, and C. Linnhoff-Popien · 2019
Later among the works it cites.
Adversarial policies: Attacking deep reinforcement learning
A. Gleave, M. Dennis, C. Wild, N. Kant, S. Levine, and S. Russell · 2019
Later among the works it cites.
Re-determinizing MCTS in Hanabi
J. Goodman · 2019
Later among the works it cites.
Recurrent independent mechanisms
A. Goyal, A. Lamb, J. Hoffmann, S. Sodhani, S. Levine, Y. Bengio, and B. Schölkopf · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Later among the works it cites.
A local approach to forward model learning: Results on the game of life game
S. M. Lucas, A. Dockhorn, V. Volz, C. Bamford, R. D. Gaina, I. Bravi, D. Perez-Liebana, S. Mostaghim, and R. Kruse · 2019
Later among the works it cites.
Efficient Evolutionary Methods for Game Agent Optimisation: Model-Based is Best
S. M. Lucas, J. Liu, I. Bravi, R. D. Gaina, J. Woodward, V. Volz, and D. Perez-Liebana · 2019
Later among the works it cites.
Evolving action abstractions for real-time planning in extensive-form games
J. R. Marino, R. O. Moraes, C. Toledo, and L. H. Lelis · 2019
Later among the works it cites.
A hybrid planning and execution approach through HTN and MCTS
X. Neufeld, S. Mostaghim, and D. Perez-Liebana · 2019
Later among the works it cites.
General Video Game AI: A Multitrack Framework for Evaluating Agents, Games, and Content Generation Algorithms
D. Perez-Liebana, J. Liu, A. Khalifa, R. D. Gaina, J. Togelius, and S. M. Lucas · 2019
Later among the works it cites.
Increasing Generality in Machine Learning through Procedural Content Generation
S. Risi and J. Togelius · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2019
Later among the works it cites.
Model-based active exploration
P. Shyam, W. Jaśkowski, and F. Gomez · 2019
Later among the works it cites.
Enhancing Rolling Horizon Evolution with Policy and Value Networks
X. Tong, W. Liu, and B. Li · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
POET: open-ended coevolution of environments and their optimized solutions
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Later among the works it cites.
Explain Your Move: Understanding Agent Actions Using Focused Feature Saliency
P. Gupta, N. Puri, S. Verma, D. Kayastha, S. Deshmukh, B. Krishnamurthy, and S. Singh · 2020
Closest in time.
Revealing Neural Network Bias to Non-Experts Through Interactive Counterfactual Examples
C. M. Myers, E. Freed, L. F. L. Pardo, A. Furqan, S. Risi, and J. Zhu · 2020
Closest in time.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Closest in time.