Fetching the paper…
Reading the bibliography…
We study stochastic policy gradient methods from the perspective of control-theoretic limitations.
Convergence of estimates under dimensionality restrictions
Lucien LeCam · 1973
Earlier work this paper cites.
Guaranteed margins for lqg regulators
John C Doyle · 1978
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
Discrete-time stochastic systems: estimation and control
Torsten Söderström · 2002
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Earlier work this paper cites.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Stephen Tu and Benjamin Recht · 2019
Earlier work this paper cites.
Learning robust control for lqr systems with multiplicative noise via policy gradient
Benjamin Gravell, Peyman Mohajerin Esfahani, and Tyler Summers · 2019
Earlier work this paper cites.
Analyzing the variance of policy gradient estimators for the linear-quadratic regulator
James A Preiss, Sébastien MR Arnold, Chen-Yu Wei, and Marius Kloft · 2019
Earlier work this paper cites.
Recovering robustness in model-free reinforcement learning
Harish K Venkataraman and Peter J Seiler · 2019
Earlier work this paper cites.
Model-free linear quadratic control via reduction to expert prediction
Yasin Abbasi-Yadkori, Nevena Lazic, and Csaba Szepesvári · 2019
Cited alongside, same era.
Learning optimal controllers for linear systems with multiplicative noise via policy gradient
Benjamin Gravell, Peyman Mohajerin Esfahani, and Tyler Summers · 2020
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2020
Cited alongside, same era.
The driver and the engineer: Reinforcement learning and robust control
Natalie Bernat, Jiexin Chen, Nikolai Matni, and John Doyle · 2020
Cited alongside, same era.
Naive exploration is optimal for online lqr
Max Simchowitz and Dylan Foster · 2020
Cited alongside, same era.
Sample complexity of linear quadratic gaussian (lqg) control for output feedback systems
Yang Zheng, Luca Furieri, Maryam Kamgarpour, and Na Li · 2021
Later among the works it cites.
Regret bounds for adaptive nonlinear control
Nicholas M Boffi, Stephen Tu, and Jean-Jacques E Slotine · 2021
Later among the works it cites.
Stabilizing dynamical systems via policy gradient methods
Juan Perdomo, Jack Umenberger, and Max Simchowitz · 2021
Later among the works it cites.
On the sample complexity of stability constrained imitation learning
Stephen Tu, Alexander Robey, Tingnan Zhang, and Nikolai Matni · 2021
Later among the works it cites.
Online policy gradient for model free learning of linear quadratic regulators with T \sqrt{T} regret
Asaf B Cassel and Tomer Koren · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anastasios Tsiamis and George J Pappas · 2021
Cited alongside, same era.
On uninformative optimal policies in adaptive lqr with unknown b-matrix
Ingvar Ziemann and Henrik Sandberg · 2021
Cited alongside, same era.
Analysis of the optimization landscape of linear quadratic gaussian (lqg) control
Yujie Tang, Yang Zheng, and Na Li · 2021
Cited alongside, same era.
On the lack of gradient domination for linear quadratic gaussian problems with incomplete state information
Hesameddin Mohammadi, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2021
Cited alongside, same era.
Global convergence of policy gradient methods to (almost) locally optimal policies
Kaiqing Zhang, Alec Koppel, Hao Zhu, and Tamer Basar
Cited in the paper.
Policy optimization for ℋ 2 \mathcal{H}_{2} linear control with ℋ ∞ \mathcal{H}_{\infty} robustness guarantee: Implicit regularization and global convergence
Kaiqing Zhang, Bin Hu, and Tamer Basar
Cited in the paper.
Ingvar Ziemann and Henrik Sandberg · 2022
Closest in time.
Learning to control linear systems can be hard
Anastasios Tsiamis, Ingvar Ziemann, Manfred Morari, Nikolai Matni, and George J Pappas · 2022
Closest in time.
Linear quadratic control using model-free reinforcement learning
Farnaz Adib Yaghmaie, Fredrik Gustafsson, and Lennart Ljung · 2022
Closest in time.
Single trajectory nonparametric learning of nonlinear dynamics
Ingvar Ziemann, Henrik Sandberg, and Nikolai Matni · 2022
Closest in time.