Fetching the paper…
Reading the bibliography…
Reinforcement learning is a powerful tool to learn the optimal policy of possibly multiple agents by interacting with the environment.
Least squares stationary optimal control and the algebraic riccati equation
Jan Willems · 1971
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Continuous-time mean-variance portfolio selection: A stochastic lq framework
Xun Yu Zhou and Duan Li · 2000
Earlier work this paper cites.
Formation constrained multi-agent control
Magnus Egerstedt and Xiaoming Hu · 2001
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Game theory and decision theory in multi-agent systems
Simon Parsons and Michael Wooldridge · 2002
Earlier work this paper cites.
Individual and mass behaviour in large population stochastic wireless power control problems: centralized and Nash equilibrium solutions
Minyi Huang, Peter E Caines, and Roland P Malhamé · 2003
Earlier work this paper cites.
Multiagent reinforcement learning for multi-robot systems: A survey
Erfu Yang and Dongbing Gu · 2004
Earlier work this paper cites.
LQ dynamic optimization and differential games
Jacob Engwerda · 2005
Earlier work this paper cites.
Some results on two-person zero-sum linear quadratic differential games
Pingjian Zhang · 2005
Earlier work this paper cites.
Learning to cooperate in multi-agent social dilemmas
Enrique Munoz de Cote, Alessandro Lazaric, and Marcello Restelli · 2006
Earlier work this paper cites.
Large population stochastic dynamic games: closed-loop mckean-vlasov systems and the Nash certainty equivalence principle
Minyi Huang, Roland P Malhamé, Peter E Caines, et al · 2006
Earlier work this paper cites.
Jeux à champ moyen. i–le cas stationnaire
Jean-Michel Lasry and Pierre-Louis Lions · 2006
Earlier work this paper cites.
Jeux à champ moyen. ii–horizon fini et contrôle optimal
Jean-Michel Lasry and Pierre-Louis Lions · 2006
Earlier work this paper cites.
Optimal control: linear quadratic methods
Brian D O Anderson and John B Moore · 2007
Earlier work this paper cites.
Mean field games
Jean-Michel Lasry and Pierre-Louis Lions · 2007
Earlier work this paper cites.
Cooperative control of distributed multi-agent systems
Jeff Shamma · 2008
Earlier work this paper cites.
Multi-agent team cooperation: A game theory approach
Elham Semsar-Kazerooni and Khashayar Khorasani · 2009
Cited alongside, same era.
Stability analysis for multi-agent systems using the incidence matrix: Quantized communication and formation control
Dimos V Dimarogonas and Karl H Johansson · 2010
Cited alongside, same era.
Optimal control in a cooperative network of smart power grids
Riccardo Minciardi and Roberto Sacile · 2011
Cited alongside, same era.
Robust mean field games with application to production of an exhaustible resource
Dario Bauso, Hamidou Tembine, and Tamer Başar · 2012
Cited alongside, same era.
Mean field games and mean field type control theory
Alain Bensoussan, Jens Frehse, Phillip Yam, et al · 2013
Cited alongside, same era.
Discrete time mean-field stochastic linear-quadratic optimal control problems
Robert Elliott, Xun Li, and Yuan-Hua Ni · 2013
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Probabilistic Theory of Mean Field Games with Applications I-II
René Carmona, François Delarue, et al · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Linear–quadratic mean-field game for stochastic delayed systems
Jianhui Huang and Na Li · 2018
Later among the works it cites.
Inequity aversion improves cooperation in intertemporal social dilemmas
Edward Hughes, Joel Z Leibo, Matthew Phillips, Karl Tuyls, Edgar Dueñez-Guzman, Antonio García Castañeda, Iain Dunning, Tina Zhu, Kevin McKee, Raphael Koster, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Risk-sensitive mean-field games
Hamidou Tembine, Quanyan Zhu, and Tamer Başar · 2013
Cited alongside, same era.
The LQR controller design of two-wheeled self-balancing robot based on the particle swarm optimization algorithm
Jian Fang · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Openai five
OpenAI · 2018
Later among the works it cites.
Markov-Nash equilibria in mean-field games with discounted cost
Naci Saldi, Tamer Basar, and Maxim Raginsky · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Global convergence of policy gradient for sequential zero-sum linear quadratic dynamic games
Jingjing Bu, Lillian J Ratliff, and Mehran Mesbahi · 2019
Later among the works it cites.
Linear-quadratic mean-field reinforcement learning: convergence of policy gradient methods
René Carmona, Mathieu Laurière, and Zongjun Tan · 2019
Later among the works it cites.
Actor-critic provably finds Nash equilibria of linear-quadratic mean-field games
Zuyue Fu, Zhuoran Yang, Yongxin Chen, and Zhaoran Wang · 2019
Later among the works it cites.
Learning mean-field games
Xin Guo, Anran Hu, Renyuan Xu, and Junzi Zhang · 2019
Later among the works it cites.
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2019
Later among the works it cites.
Approximate Nash equilibria in partially observed stochastic games with mean-field interactions
Naci Saldi, Tamer Başar, and Maxim Raginsky · 2019
Later among the works it cites.
Alphastar: Mastering the real-time strategy game starcraft ii
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, et al · 2019
Later among the works it cites.
Policy optimization provably converges to Nash equilibria in zero-sum linear quadratic games
Kaiqing Zhang, Zhuoran Yang, and Tamer Basar · 2019
Later among the works it cites.
On the convergence of model free learning in mean field games
Romuald Elie, Julien Perolat, Mathieu Laurière, Matthieu Geist, and Olivier Pietquin · 2020
Closest in time.