Fetching the paper…
Reading the bibliography…
We consider scenarios from the real-time strategy game StarCraft as new benchmarks for reinforcement learning algorithms.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
1933
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
1952
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
1988
Earlier work this paper cites.
Q-learning
1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
1995
Earlier work this paper cites.
A one-measurement form of simultaneous perturbation stochastic approximation
1997
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm
1998
Earlier work this paper cites.
Reinforcement learning: An introduction
1998
Earlier work this paper cites.
Team-partitioned, opaque-transition reinforcement learning
1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
1999
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
2002
Earlier work this paper cites.
The cross entropy method for fast policy search
2003
Cited alongside, same era.
Extending q-learning to general adaptive multi-agent systems
2003
Cited alongside, same era.
Exploration exploitation in go: Uct for monte-carlo go
2006
Cited alongside, same era.
Hierarchical multi-agent reinforcement learning
2006
Cited alongside, same era.
Learning tetris using the noisy cross-entropy method
2006
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
2008
Cited alongside, same era.
Policy gradients with parameter-based exploration for control
2008
Optimal rates for zero-order convex optimization: the power of two function evaluations
2013
Later among the works it cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
2013
Later among the works it cites.
Playing atari with deep reinforcement learning
2013
Later among the works it cites.
A survey of real-time strategy game ai research and competition in starcraft
2013
Later among the works it cites.
Evolving effective micro behaviors in rts game
2014
Later among the works it cites.
Deterministic policy gradient algorithms
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Structured prediction with reinforcement learning
2009
Cited alongside, same era.
Exploring parameter space in reinforcement learning
2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
2011
Cited alongside, same era.
A bayesian model for rts units control applied to starcraft
2011
Cited alongside, same era.
Fast heuristic search for rts game combat scenarios
2012
Cited alongside, same era.
2015
Later among the works it cites.
Human-level control through deep reinforcement learning
2015
Later among the works it cites.
Deep reinforcement learning with double q-learning
2015
Later among the works it cites.
Exploratory gradient boosting for reinforcement learning in complex domains
2016
Closest in time.
Control of memory, active perception, and action in minecraft
2016
Closest in time.
Deep exploration via bootstrapped dqn
2016
Closest in time.
Generalization and exploration via randomized value functions
2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
2016
Closest in time.
Learning multiagent communication with backpropagation
2016
Closest in time.