Fetching the paper…
Reading the bibliography…
Multiagent systems where agents interact among themselves and with a stochastic environment can be formalized as stochastic games.
Calculus: Multi-variable Calculus and Linear Algebra, with Applications to Differential Equations and Probability
T. Apostol · 1969
Earlier work this paper cites.
Noncooperative and dominant player solutions in discrete dynamic games
F. Kydland · 1975
Earlier work this paper cites.
Optimum Systems Control
A. P. Sage and C. C. White · 1977
Earlier work this paper cites.
The great fish war: An example using a dynamic Cournot-Nash solution
D. Levhari and L. J. Mirman · 1980
Earlier work this paper cites.
Open-loop and closed-loop equilibria in dynamic games with many players
D. Fudenberg and D. K. Levine · 1988
Earlier work this paper cites.
Dynamic Noncooperative Game Theory
T. Basar and G. J. Olsder · 1999
Earlier work this paper cites.
On actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2003
Earlier work this paper cites.
Dynamic Programming and Optimal Control , volume 2
D. P. Bertsekas · 2007
Earlier work this paper cites.
Fitted natural actor-critic: A new algorithm for continuous state-action MDPs
F. S. Melo and M. Lopes · 2008
Cited alongside, same era.
Generalized nash equilibrium problems
F. Facchinei and C. Kanzow · 2010
Cited alongside, same era.
A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications
W. B. Powell and J. Ma · 2011
Cited alongside, same era.
Reinforcement learning in continuous state and action spaces
H. Van Hasselt · 2012
Cited alongside, same era.
Discrete–Time Stochastic Control and Dynamic Potential Games: The Euler–Equation Approach
D. González-Sánchez and O. Hernández-Lerma · 2013
Cited alongside, same era.
CVX: Matlab software for disciplined convex programming, version 2.1
Two-timescale algorithms for learning nash equilibria in general-sum stochastic games
H. Prasad, P. LA, and S. Bhatnagar · 2015
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
Learning in constrained stochastic dynamic potential games
S. Valcarcel Macua, S. Zazo, and J. Zazo · 2016
Later among the works it cites.
Counterfactual multi-agent policy gradients
J. N. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2017
Later among the works it cites.
Learning nash equilibrium for general-sum markov games from batch data
J. Pérolat, F. Strub, B. Piot, and O. Pietquin · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Grant and S. Boyd · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Dynamic potential games with constraints: Fundamentals and applications in communications
S. Zazo, S. V. Macua, M. Sánchez-Fernández, and J. Zazo
Cited in the paper.
Dynamic potential games with constraints: Fundamentals and applications in communications
S. Zazo, S. Valcarcel Macua, M. Sánchez-Fernández, and J. Zazo
Cited in the paper.
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. F. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, and T. Graepel · 2017
Later among the works it cites.