Multiagent reinforcement learning: theoretical framework and an algorithm
Hu, J., Wellman, M. P., et al. (1998) · 1998
Cited alongside, same era.
Multiagent reinforcement learning in stochastic games
Hu, J. and Wellman, M. P. (1999) · 1999
Cited alongside, same era.
Finite-sample convergence rates for q-learning and indirect algorithms
Kearns, M. J. and Singh, S. P. (1999) · 1999
Cited alongside, same era.
On the complexity of policy iteration
Mansour, Y. and Singh, S. (1999) · 1999
Cited alongside, same era.
An analysis of stochastic game theory for multiagent reinforcement learning
Bowling, M. and Veloso, M. (2000) · 2000
Cited alongside, same era.
Rational and convergent learning in stochastic games
Bowling, M. and Veloso, M. (2001) · 2001
Cited alongside, same era.
Markov perfect equilibrium. I. Observable actions
Maskin, E. and Tirole, J. (2001) · 2001
Cited alongside, same era.
Nash q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P. (2003) · 2003
Cited alongside, same era.
On the sample complexity of reinforcement learning
Kakade, S. M. (2003) · 2003
Cited alongside, same era.
A new complexity result on solving the Markov decision problem
Ye, Y. (2005) · 2005
Cited alongside, same era.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Even-Dar, E., Mannor, S., and Mansour, Y. (2006) · 2006
Cited alongside, same era.
Reinforcement learning with a near optimal rate of convergence
Azar, M. G., Munos, R., Ghavamzadeh, M., and Kappen, H. (2011a)
Cited in the paper.