Fetching the paper…
Reading the bibliography…
First order policy optimization has been widely used in reinforcement learning.
R. E. Kalman, “Contributions to the theory of optimal control,” Bol. soc. mat. mexicana , vol. 5, no. 2, pp. 102–119, 1960
1960
Earlier work this paper cites.
G. Hewer, “An iterative technique for the computation of the steady state gains for the discrete optimal regulator,” IEEE Trans. Automat. Contr. , vol. 16, no. 4, pp. 382–384, 1971
1971
Earlier work this paper cites.
1978
Earlier work this paper cites.
P. Lancaster and L. Rodman, Algebraic riccati equations . Clarendon press, 1995
1995
Earlier work this paper cites.
K. Zhou, J. C. Doyle, K. Glover, et al. , Robust and optimal control . Prentice hall New Jersey, 1996, vol. 40
1996
Earlier work this paper cites.
V. Balakrishnan and L. Vandenberghe, “Semidefinite programming duality and linear time-invariant systems,” IEEE Trans. Automat. Contr. , vol. 48, no. 1, pp. 30–41, 2003
2003
Earlier work this paper cites.
D. Bertsekas, Dynamic programming and optimal control: Volume I . Athena scientific, 2012, vol. 1
2012
Earlier work this paper cites.
G. E. Dullerud and F. Paganini, A course in robust control theory: a convex approach . Springer Science & Business Media, 2013, vol. 36
2013
Earlier work this paper cites.
R. Ge, F. Huang, C. Jin, and Y. Yuan, “Escaping from saddle points—online stochastic gradient for tensor decomposition,” in Conference on learning theory . PMLR, 2015, pp. 797–842
2015
Cited alongside, same era.
2016
Cited alongside, same era.
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan, “How to escape saddle points efficiently,” in International Conference on Machine Learning . PMLR, 2017, pp. 1724–1732
2017
Cited alongside, same era.
2018
Cited alongside, same era.
K. Zhang, Z. Yang, and T. Basar, “Policy optimization provably converges to nash equilibria in zero-sum linear quadratic games,” Advances in Neural Information Processing Systems , vol. 32, pp. 11 602–11 614, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
L. Furieri, Y. Zheng, and M. Kamgarpour, “Learning the globally optimal distributed LQ regulator,” in Learning for Dynamics and Control . PMLR, 2020, pp. 287–297
2020
Later among the works it cites.
Y. Sun and M. Fazel, “Learning optimal controllers by policy gradient: Global optimality via convex parameterization,” in 2021 60th IEEE Conference on Decision and Control (CDC) . IEEE, 2021, pp. 4576–4581
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Recht, “A tour of reinforcement learning: The view from continuous control,” Annual Review of Control, Robotics, and Autonomous Systems , vol. 2, pp. 253–279, 2019
2019
Cited alongside, same era.
H. Mohammadi, A. Zare, M. Soltanolkotabi, and M. R. Jovanović, “Global exponential convergence of gradient methods over the nonconvex landscape of the linear quadratic regulator,” in 2019 IEEE 58th Conference on Decision and Control (CDC) . IEEE, 2019, pp. 7474–7479
2019
Cited alongside, same era.
D. Malik, A. Pananjady, K. Bhatia, K. Khamaru, P. Bartlett, and M. Wainwright, “Derivative-free methods for policy optimization: Guarantees for linear quadratic systems,” in The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 2019, pp. 2916–2925
2019
Cited alongside, same era.
S. Tu and B. Recht, “The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint,” in Conference on Learning Theory . PMLR, 2019, pp. 3036–3083
2019
Cited alongside, same era.
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.