Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning has been successfully applied to a number of challenging problems.
The bias and moment matrix of the general k-class estimators of the parameters in simultaneous equations
Anirudh L Nagar · 1959
Earlier work this paper cites.
A counterexample in stochastic optimum control
Hans S Witsenhausen · 1968
Earlier work this paper cites.
Separation of estimation and control for discrete time systems
Hans S Witsenhausen · 1971
Earlier work this paper cites.
Team decision theory and information structures in optimal control problems–part i
Yu-Chi Ho et al · 1972
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Reinforcement learning applied to linear quadratic regulation
Steven J Bradtke · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Robust and optimal control , volume 40
Kemin Zhou, John Comstock Doyle, and Keith Glover · 1996
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay R. Konda and John N. Tsitsiklis · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal nash equilibrium in team markov games
Xiaofeng Wang and Tuomas Sandholm · 2003
Earlier work this paper cites.
Multiagent reinforcement learning for multi-robot systems: A survey
Erfu Yang and Dongbing Gu · 2004
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Liviu Panait and Sean Luke · 2005
Earlier work this paper cites.
Optimal control: linear quadratic methods
Brian DO Anderson and John B Moore · 2007
Earlier work this paper cites.
Incremental natural actor-critic algorithms
Shalabh Bhatnagar, Mohammad Ghavamzadeh, Mark Lee, and Richard S Sutton · 2008
Earlier work this paper cites.
Distributed lqr design for identical dynamically decoupled systems
Francesco Borrelli and TamÁs Keviczky · 2008
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Bu, Robert Babu, Bart De Schutter, et al · 2008
Earlier work this paper cites.
Distributed control for identical dynamically coupled systems: A decomposition approach
Paolo Massioni and Michel Verhaegen · 2008
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Reinforcement learning and adaptive dynamic programming for feedback control
Frank L Lewis and Draguna Vrabie · 2009
Earlier work this paper cites.
Controllability of multi-agent systems from a graph-theoretic perspective
Amirreza Rahmani, Meng Ji, Mehran Mesbahi, and Magnus Egerstedt · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Cited alongside, same era.
An actor–critic algorithm with function approximation for discounted cost constrained markov decision processes
Shalabh Bhatnagar · 2010
Cited alongside, same era.
A convergent online single time scale actor critic algorithm
Dotan Di Castro and Ron Meir · 2010
Cited alongside, same era.
Optimal control in a cooperative network of smart power grids
Riccardo Minciardi and Roberto Sacile · 2011
Cited alongside, same era.
Dynamic programming and optimal control, Vol. II, 4th Edition
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P How, and John Vian · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2017
Later among the works it cites.
Learning deep mean field games for modeling large population behavior
Jiachen Yang, Xiaojing Ye, Rakshit Trivedi, Huan Xu, and Hongyuan Zha · 2017
Later among the works it cites.
Flight control law of unmanned aerial vehicles baped on robust servo linear quadratic regulator and kalman filtering
Yongfeng Zhi, Gaoshang Li, Qun Song, Ke Yu, and Jun Zhang · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dimitri P Bertsekas · 2012
Cited alongside, same era.
Sub-optimal distributed control law with h2 performance for identical dynamically coupled linear systems
Paresh Deshpande, PP Menon, Christopher Edwards, and Ian Postlethwaite · 2012
Cited alongside, same era.
Multi-agent reinforcement learning for integrated network of adaptive traffic signal controllers (marlin-atsc)
Samah El-Tantawy and Baher Abdulhai · 2012
Cited alongside, same era.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
Ivo Grondman, Lucian Busoniu, Gabriel AD Lopes, and Robert Babuska · 2012
Cited alongside, same era.
Social optima in mean field lqg control: centralized and decentralized strategies
Minyi Huang, Peter E Caines, and Roland P Malhamé · 2012
Cited alongside, same era.
Control of mckean–vlasov dynamics versus mean field games
René Carmona, François Delarue, and Aimé Lachapelle · 2013
Cited alongside, same era.
Discrete time mean-field stochastic linear-quadratic optimal control problems
Robert Elliott, Xun Li, and Yuan-Hua Ni · 2013
Cited alongside, same era.
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2018
Later among the works it cites.
Primal-dual algorithm for distributed reinforcement learning: distributed gtd
Donghwan Lee, Hyungjin Yoon, and Naira Hovakimyan · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter L Bartlett, and Martin J Wainwright · 2018
Later among the works it cites.
Deep reinforcement learning for event-driven multi-agent decision processes
Kunal Menda, Yi-Chun Chen, Justin Grana, James W. Bono, Brendan D. Tracey, Mykel J. Kochenderfer, and David H. Wolpert · 2018
Later among the works it cites.
Credit assignment for collective multiagent rl with global rewards
Duc Thien Nguyen, Akshat Kumar, and Hoong Chuin Lau · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
A tour of reinforcement learning: The view from continuous control
Benjamin Recht · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Later among the works it cites.
Stephen Tu and Benjamin Recht · 2018
Later among the works it cites.
Distributed lqr methods for networks of non-identical plants
Eleftherios E. Vlahakis and George D. Halikias · 2018
Later among the works it cites.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Later among the works it cites.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Başar · 2018
Later among the works it cites.
Distributed q-learning for dynamically decoupled systems
Siavash Alemzadeh and Mehran Mesbahi · 2019
Closest in time.
Lqr through the lens of first order methods: Discrete-time case
Jingjing Bu, Afshin Mesbahi, Maryam Fazel, and Mehran Mesbahi · 2019
Closest in time.
A tour of reinforcement learning: The view from continuous control
Benjamin Recht · 2019
Closest in time.
Linear quadratic regulator controller (lqr) for ar. drone’s safe landing
Gembong Edhi Setyawan, Wijaya Kurniawan, and Amroy Casro Lumban Gaol · 2019
Closest in time.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M. Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, Timo Ewalds, Dan Horgan, Manuel Kroiss, Ivo Danihelka, John Agapiou, Junhyuk Oh, Valentin Dalibard, David Choi, Laurent Sifre, Yury Sulsky, Sasha Vezhnevets, James Molloy, Trevor Cai, David Budden, Tom Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Toby Pohlen, Yuhuai Wu, Dani Yogatama, Julia Cohen, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Chris Apps, Koray Kavukcuoglu, Demis Hassabis, and David Silver · 2019
Closest in time.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Zhuoran Yang, Yongxin Chen, Mingyi Hong, and Zhaoran Wang · 2019
Closest in time.
Distributed control design for heterogeneous interconnected systems
Yvonne R Sturz, Annika Eichler, and Roy S Smith · 2020
Closest in time.
Trajectory optimization for uav-to-device underlaid cellular networks by mean-field-type control
Hongliang Zhang, Zhu Han, and H. Vincent Poor · 2020
Closest in time.
Stabilizing dynamical systems via policy gradient methods
Juan Perdomo, Jack Umenberger, and Max Simchowitz · 2021
Closest in time.