Fetching the paper…
Reading the bibliography…
This paper considers a distributed reinforcement learning problem for decentralized linear quadratic control with partial state observations and local costs.
1912
Earlier work this paper cites.
H. S. Witsenhausen, “A counterexample in stochastic optimum control,” SIAM Journal on Control , vol. 6, no. 1, pp. 131–147, 1968
1968
Earlier work this paper cites.
P. M. Gahinet, A. J. Laub, C. S. Kenney, and G. A. Hewer, “Sensitivity of the stable discrete-time Lyapunov equation,” IEEE Transactions on Automatic Control , vol. 35, no. 11, pp. 1209–1217, 1990
1990
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3-4, pp. 229–256, 1992
1992
Earlier work this paper cites.
Y. U. Cao, A. S. Fukunaga, and A. B. Kahng, “Cooperative mobile robotics: Antecedents and directions,” Autonomous Robots , vol. 4, no. 1, pp. 7–27, 1997
1997
Earlier work this paper cites.
L. Peshkin, K.-E. Kim, N. Meuleau, and L. P. Kaelbling, “Learning to cooperate via policy search,” in Proceedings of the Sixteenth Conference on Uncertainty in Artificial Intelligence , 2000, pp. 489–496
2000
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems . MIT Press, 2000, vol. 12, pp. 1057–1063
2000
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, “The complexity of decentralized control of Markov decision processes,” Mathematics of Operations Research , vol. 27, no. 4, pp. 819–840, 2002
2002
Earlier work this paper cites.
K. B. Ariyur and M. Krstic, Real-Time Optimization by Extremum-Seeking Control . John Wiley & Sons, 2003
2003
Earlier work this paper cites.
G. A. F. Seber and A. J. Lee, Linear Regression Analysis , 2nd ed. John Wiley & Sons, 2003
2003
Earlier work this paper cites.
L. Xiao and S. Boyd, “Fast linear iterations for distributed averaging,” Systems & Control Letters , vol. 53, no. 1, pp. 65–78, 2004
2004
Earlier work this paper cites.
Y.-C. Wang and J. M. Usher, “Application of reinforcement learning for agent-based production scheduling,” Engineering Applications of Artificial Intelligence , vol. 18, no. 1, pp. 73–82, 2005
2005
Earlier work this paper cites.
M. Rotkowitz and S. Lall, “A characterization of convex problems in decentralized control,” IEEE Transactions on Automatic Control , vol. 50, no. 12, pp. 1984–1996, 2005
2005
Earlier work this paper cites.
A. D. Flaxman, A. T. Kalai, A. T. Kalai, and H. B. McMahan, “Online convex optimization in the bandit setting: Gradient descent without a gradient,” in Proceedings of the Sixteenth Annual ACM-SIAM Symposium on Discrete Algorithms , 2005, pp. 385–394
2005
Earlier work this paper cites.
L. Bakule, “Decentralized control: An overview,” Annual Reviews in Control , vol. 32, no. 1, pp. 87–98, 2008
2008
Earlier work this paper cites.
K. J. Åström and B. Wittenmark, Adaptive Control , 2nd ed. Dover Publications, 2008
2008
Earlier work this paper cites.
M. Riedmiller, T. Gabel, R. Hafner, and S. Lange, “Reinforcement learning for robot soccer,” Autonomous Robots , vol. 27, no. 1, pp. 55–73, 2009
2009
Cited alongside, same era.
A. L. C. Bazzan, “Opportunities for multiagent systems and multiagent reinforcement learning in traffic control,” Autonomous Agents and Multi-Agent Systems , vol. 18, no. 3, pp. 342–375, 2009
2009
Cited alongside, same era.
M. Pipattanasomporn, H. Feroze, and S. Rahman, “Multi-agent systems in a distributed smart grid: Design and implementation,” in 2009 IEEE/PES Power Systems Conference and Exposition , 2009, pp. 1–8
2009
Cited alongside, same era.
K. Mårtensson and A. Rantzer, “Gradient methods for iterative distributed control synthesis,” in Proceedings of the 48h IEEE Conference on Decision and Control (CDC) held jointly with 2009 28th Chinese Control Conference . IEEE, 2009, pp. 549–554
2009
Cited alongside, same era.
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian, “Deep decentralized multi-task multi-agent reinforcement learning under partial observability,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 70, 2017, pp. 2681–2690
2017
Later among the works it cites.
G. Qu and N. Li, “Harnessing smoothness to accelerate distributed optimization,” IEEE Transactions on Control of Network Systems , vol. 5, no. 3, pp. 1245–1260, 2017
2017
Later among the works it cites.
S. Shah, D. Dey, C. Lovett, and A. Kapoor, “Airsim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and Service Robotics , ser. Springer Proceedings in Advanced Robotics, M. Hutter and R. Siegwart, Eds. Springer International Publishing, 2018, vol. 5, pp. 621–635
2018
Later among the works it cites.
M. Fazel, R. Ge, S. Kakade, and M. Mesbahi, “Global convergence of policy gradient methods for the linear quadratic regulator,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80, 2018, pp. 1467–1476
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Al Alam, A. Gattami, and K. H. Johansson, “Suboptimal decentralized controller design for chain structures: Applications to vehicle formations,” in Proceedings of the 50th IEEE Conference on Decision and Control (CDC) and European Control Conference , 2011, pp. 6894–6900
2011
Cited alongside, same era.
F. L. Lewis, D. L. Vrabie, and V. L. Syrmos, Optimal Control , 3rd ed. John Wiley & Sons, 2012
2012
Cited alongside, same era.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Cited alongside, same era.
O. Shamir, “On the complexity of bandit and derivative-free stochastic convex optimization,” in Proceedings of the 26th Annual Conference on Learning Theory , ser. Proceedings of Machine Learning Research, vol. 30, 2013, pp. 3–24
2013
Cited alongside, same era.
M. I. Abouheaf, F. L. Lewis, K. G. Vamvoudakis, S. Haesaert, and R. Babuska, “Multi-agent discrete-time graphical games and reinforcement learning solutions,” Automatica , vol. 50, no. 12, pp. 3038–3053, 2014
2014
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 32, 2014, pp. 387–395
2014
Cited alongside, same era.
X. Zhang, W. Shi, X. Li, B. Yan, A. Malkawi, and N. Li, “Decentralized temperature control via HVAC systems in energy efficient buildings: An approximate solution procedure,” in Proceedings of 2016 IEEE Global Conference on Signal and Information Processing , 2016, pp. 936–940
2016
Cited alongside, same era.
H. Zhang, H. Jiang, Y. Luo, and G. Xiao, “Data-driven optimal consensus control for discrete-time multi-agent systems with unknown dynamics using reinforcement learning method,” IEEE Transactions on Industrial Electronics , vol. 64, no. 5, pp. 4091–4100, 2016
2016
Cited alongside, same era.
2018
Later among the works it cites.
M. Gagrani and A. Nayyar, “Thompson sampling for some decentralized control problems,” in Proceedings of the 57th IEEE Conference on Decision and Control (CDC) , 2018, pp. 1053–1058
2018
Later among the works it cites.
J. N. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson, “Counterfactual multi-agent policy gradients,” in The Thirty-Second AAAI Conference on Artificial Intelligence , 2018, pp. 2974–2982
2018
Later among the works it cites.
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Basar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80, 2018, pp. 5872–5881
2018
Later among the works it cites.
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu, “On the sample complexity of the linear quadratic regulator,” Foundations of Computational Mathematics , pp. 1–47, 2019
2019
Closest in time.
Z. Yang, Y. Chen, M. Hong, and Z. Wang, “Provably global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost,” in Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019, vol. 32, pp. 8351–8363
2019
Closest in time.
H. Mania, S. Tu, and B. Recht, “Certainty equivalence is efficient for linear quadratic control,” in Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019, vol. 32, pp. 10 154–10 164
2019
Closest in time.
S. Oymak and N. Ozay, “Non-asymptotic identification of LTI systems from a single trajectory,” in 2019 American Control Conference (ACC) , 2019, pp. 5655–5661
2019
Closest in time.
2019
Closest in time.
K. Zhang, E. Miehling, and T. Başar, “Online planning for decentralized stochastic control with partial history sharing,” in 2019 American Control Conference (ACC) . IEEE, 2019, pp. 3544–3550
2019
Closest in time.
H. Feng and J. Lavaei, “On the exponential number of connected components for the feasible set of optimal decentralized control problems,” in 2019 American Control Conference , 2019, pp. 1430–1437
2019
Closest in time.
D. Malik, A. Pananjady, K. Bhatia, K. Khamaru, P. L. Bartlett, and M. J. Wainwright, “Derivative-free methods for policy optimization: Guarantees for linear quadratic systems,” Journal of Machine Learning Research , vol. 21, no. 21, pp. 1–51, 2020
2020
Closest in time.