Fetching the paper…
Reading the bibliography…
We consider a subclass of $n$-player stochastic games, in which players have their own internal state/action spaces while they are coupled through their payoff functions.
J. F. Nash, “Equilibrium points in n n -person games,” Proceedings of the National Academy of Sciences , vol. 36, no. 1, pp. 48–49, 1950
1950
Earlier work this paper cites.
L. S. Shapley, “Stochastic games,” Proceedings of the National Academy of Sciences , vol. 39, no. 10, pp. 1095–1100, 1953
1953
Earlier work this paper cites.
J. B. Rosen, “Existence and uniqueness of equilibrium points for concave n n -person games,” Econometrica: Journal of the Econometric Society , pp. 520–534, 1965
1965
Earlier work this paper cites.
P. Hall and C. C. Heyde, Martingale Limit Theory and its Application . Academic Press, 1980
1980
Earlier work this paper cites.
A. S. Nemirovski and D. B. Yudin, Problem Complexity and Method Efficiency in Optimization . Wiley-Interscience Series in Discrete Mathematics, Wiley, New York, 1983
1983
Earlier work this paper cites.
R. J. Aumann, “Correlated equilibrium as an expression of Bayesian rationality,” Econometrica: Journal of the Econometric Society , pp. 1–18, 1987
1987
Earlier work this paper cites.
D. Monderer and L. S. Shapley, “Potential games,” Games and Economic Behavior , vol. 14, no. 1, pp. 124–143, 1996
1996
Earlier work this paper cites.
T. Başar and G. J. Olsder, Dynamic Noncooperative Game Theory . 2nd Ed, SIAM, 1999
1999
Earlier work this paper cites.
E. Altman, Constrained Markov Decision Processes . CRC Press, 1999
1999
Earlier work this paper cites.
A. Beck and M. Teboulle, “Mirror descent and nonlinear projected subgradient methods for convex optimization,” Operations Research Letters , vol. 31, no. 3, pp. 167–175, 2003
2003
Earlier work this paper cites.
S. Boyd, S. P. Boyd, and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
E. Altman, K. Avrachenkov, R. Marquez, and G. Miller, “Zero-sum constrained stochastic games with independent state processes,” Mathematical Methods of Operations Research , vol. 62, no. 3, pp. 375–386, 2005
2005
Earlier work this paper cites.
N. Cesa-Bianchi and G. Lugosi, Prediction, Learning, and Games . Cambridge University Press, 2006
2006
Earlier work this paper cites.
E. Altman, K. Avratchenkov, N. Bonneau, M. Debbah, R. El-Azouzi, and D. S. Menasché, “Constrained stochastic games in wireless networks,” in IEEE GLOBECOM 2007-IEEE Global Telecommunications Conference . IEEE, 2007, pp. 315–320
2007
Earlier work this paper cites.
E. Altman, K. Avrachenkov, N. Bonneau, M. Debbah, R. El-Azouzi, and D. S. Menasche, “Constrained cost-coupled stochastic games with independent state processes,” Operations Research Letters , vol. 36, no. 2, pp. 160–164, 2008
2008
Earlier work this paper cites.
C. Daskalakis, P. W. Goldberg, and C. H. Papadimitriou, “The complexity of computing a Nash equilibrium,” SIAM Journal on Computing , vol. 39, no. 1, pp. 195–259, 2009
2009
Earlier work this paper cites.
E. Even-Dar, Y. Mansour, and U. Nadav, “On the convergence of regret minimization dynamics in concave games,” in Proceedings of the Forty-First Annual ACM Symposium on Theory of Computing , 2009, pp. 523–532
2009
Earlier work this paper cites.
E. Even-Dar, S. M. Kakade, and Y. Mansour, “Online Markov decision processes,” Mathematics of Operations Research , vol. 34, no. 3, pp. 726–736, 2009
2009
Earlier work this paper cites.
M. O. Jackson, Social and Economic Networks . Princeton University Press, 2010
2010
Earlier work this paper cites.
T. Alpcan and T. Başar, Network Security: A Decision and Game-theoretic Approach . Cambridge University Press, 2010
2010
Cited alongside, same era.
C. Daskalakis, “On the complexity of approximating a Nash equilibrium,” ACM Transactions on Algorithms (TALG) , vol. 9, no. 3, pp. 1–35, 2013
2013
Cited alongside, same era.
G. Neu, A. György, C. Szepesvari, and A. Antos, “Online Markov decision processes under bandit feedback,” IEEE Transactions on Automatic Control , vol. 59, no. 3, pp. 676–691, 2013
2013
Cited alongside, same era.
S. Bubeck, “Convex optimization: Algorithms and complexity,” arXiv preprint arXiv:1405.4980 , 2014
2014
Cited alongside, same era.
T. Roughgarden, Twenty Lectures on Algorithmic Game Theory . Cambridge University Press, 2016
2016
Cited alongside, same era.
C. Daskalakis, D. J. Foster, and N. Golowich, “Independent policy gradient methods for competitive reinforcement learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 5527–5540, 2020
2020
Later among the works it cites.
Y. Jin and A. Sidford, “Efficiently solving MDPs with stochastic mirror descent,” in International Conference on Machine Learning . PMLR, 2020, pp. 4890–4900
2020
Later among the works it cites.
M. A. uz Zaman, K. Zhang, E. Miehling, and T. Bașar, “Reinforcement learning in non-stationary discrete-time linear-quadratic mean-field games,” in 2020 59th IEEE Conference on Decision and Control (CDC) . IEEE, 2020, pp. 2278–2284
2020
Later among the works it cites.
B. Gao and L. Pavel, “Continuous-time discounted mirror descent dynamics in monotone concave games,” IEEE Transactions on Automatic Control , vol. 66, no. 11, pp. 5451–5458, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Z. Zhou, P. Mertikopoulos, A. L. Moustakas, N. Bambos, and P. Glynn, “Mirror descent learning in continuous games,” in 2017 IEEE 56th Annual Conference on Decision and Control (CDC) . IEEE, 2017, pp. 5776–5783
2017
Cited alongside, same era.
D. A. Levin and Y. Peres, Markov Chains and Mixing Times . American Mathematical Soc., 2017
2017
Cited alongside, same era.
T. Başar and G. Zaccour, Handbook of Dynamic Game Theory . Springer, 2018
2018
Cited alongside, same era.
S. R. Etesami, W. Saad, N. B. Mandayam, and H. V. Poor, “Stochastic games for the smart grid energy management with prospect prosumers,” IEEE Transactions on Automatic Control , vol. 63, no. 8, pp. 2327–2342, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Golowich, S. Pattathil, and C. Daskalakis, “Tight last-iterate convergence rates for no-regret learning in multi-player games,” Advances in Neural Information Processing Systems , vol. 33, pp. 20 766–20 778, 2020
2020
Later among the works it cites.
K. Zhang, Z. Yang, and T. Başar, “Multi-agent reinforcement learning: A selective overview of theories and algorithms,” Handbook of Reinforcement Learning and Control, Springer , pp. 321–384, 2021
2021
Later among the works it cites.
S. Qiu, X. Wei, J. Ye, Z. Wang, and Z. Yang, “Provably efficient fictitious play policy optimization for zero-sum Markov games with structured transitions,” in International Conference on Machine Learning . PMLR, 2021, pp. 8715–8725
2021
Later among the works it cites.
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan, “On the theory of policy gradient methods: Optimality, approximation, and distribution shift,” Journal of Machine Learning Research , vol. 22, no. 98, pp. 1–76, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Tian, Y. Wang, T. Yu, and S. Sra, “Online learning in unknown Markov games,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 279–10 288
2021
Later among the works it cites.
M. Sayin, K. Zhang, D. Leslie, T. Başar, and A. Ozdaglar, “Decentralized Q-learning in zero-sum Markov games,” Advances in Neural Information Processing Systems , vol. 34, pp. 18 320–18 334, 2021
2021
Later among the works it cites.
B. M. Hambly, R. Xu, and H. Yang, “Policy gradient methods find the Nash equilibrium in N-player general-sum linear-quadratic games,” Preprint, submitted August 2, https://dx.doi.org/10.2139/ssrn.3894471 , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
M. O. Sayin, F. Parise, and A. Ozdaglar, “Fictitious play in zero-sum stochastic games,” SIAM Journal on Control and Optimization , vol. 60, no. 4, pp. 2095–2114, 2022
2022
Closest in time.