Fetching the paper…
Reading the bibliography…
We consider a general-sum N-player linear-quadratic game with stochastic dynamics over a finite horizon and prove the global convergence of the natural policy gradient method to the Nash equilibrium.
Trace bounds on the solution of the algebraic matrix riccati and lyapunov equation
Sheng-De Wang, Te-Son Kuo, and Chen-Fa Hsu · 1986
Earlier work this paper cites.
A matrix inequality associated with bounds on solutions of algebraic riccati and lyapunov equations
J. Saniuk and I. Rhodes · 1987
Earlier work this paper cites.
Dynamic Non-Cooperative Game Theory
Tamer Başar and Geert Jan Olsder · 1998
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder Singh, Michael J Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Large population stochastic dynamic games: closed-loop mckean-vlasov systems and the nash certainty equivalence principle
Minyi Huang, Roland P Malhamé, and Peter E Caines · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Mean field games
Jean-Michel Lasry and Pierre-Louis Lions · 2007
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor Lesser · 2010
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Cited alongside, same era.
Policy-gradient algorithms have no guarantees of convergence in linear quadratic games
Eric Mazumdar, Lillian J Ratliff, Michael I Jordan, and S Shankar Sastry · 2019
Later among the works it cites.
Convergence of multi-agent learning with a finite step size in general-sum games
Xinliang Song, Tonghan Wang, and Chongjie Zhang · 2019
Later among the works it cites.
Policy optimization provably converges to Nash equilibria in zero-sum linear quadratic games
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2019
Later among the works it cites.
Implicit learning dynamics in stackelberg games: Equilibria characterization, convergence analysis, and empirical study
Tanner Fiez, Benjamin Chasnov, and Lillian Ratliff · 2020
Later among the works it cites.
Policy gradient methods for the noisy linear quadratic regulator over a finite horizon
Ben M Hambly, Renyuan Xu, and Huining Yang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The mechanics of n-player differentiable games
David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
Real-time bidding with multi-agent reinforcement learning in display advertising
Junqi Jin, Chengru Song, Han Li, Kun Gai, Jun Wang, and Weinan Zhang · 2018
Cited alongside, same era.
Global convergence of policy gradient for sequential zero-sum linear quadratic dynamic games
Jingjing Bu, Lillian J Ratliff, and Mehran Mesbahi · 2019
Cited alongside, same era.
Differentiable game mechanics
Alistair Letcher, David Balduzzi, Sébastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel · 2019
Cited alongside, same era.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Cited alongside, same era.
Later among the works it cites.
On gradient-based learning in continuous games
Eric Mazumdar, Lillian J Ratliff, and S Shankar Sastry · 2020
Later among the works it cites.
Reinforcement learning in nonzero-sum linear quadratic deep structured games: Global convergence of policy optimization
Masoud Roudneshin, Jalal Arabneydi, and Amir G Aghdam · 2020
Later among the works it cites.
Reinforcement learning in continuous time and space: A stochastic control approach
Haoran Wang, Thaleia Zariphopoulou, and Xun Yu Zhou · 2020
Later among the works it cites.
Sample efficient reinforcement learning with reinforce
Junzi Zhang, Jongho Kim, Brendan O’Donoghue, and Stephen Boyd · 2020
Later among the works it cites.
Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon
Matteo Basei, Xin Guo, Anran Hu, and Yufei Zhang · 2021
Closest in time.
Kaiqing Zhang, Xiangyuan Zhang, Bin Hu, and Tamer Başar · 2021
Closest in time.