Fetching the paper…
Reading the bibliography…
We introduce the first direct policy search algorithm which provably converges to the globally optimal $\textit{dynamic}$ filter for the classical problem of predicting the outputs of a linear dynamical system, given noisy, partial observations.
A New Approach to Linear Filtering and Prediction Problems
R. E. Kalman · 1960
Earlier work this paper cites.
The importance of Kalman filtering methods for economic systems
Michael Athans · 1974
Earlier work this paper cites.
Stability of discrete linear feedback systems
Vladimír Kučera · 1975
Earlier work this paper cites.
Some geometric questions in the theory of linear systems
Roger Brockett · 1976
Earlier work this paper cites.
On guaranteed stability of uncertain linear systems via linear control
JS Thorp and BR Barmish · 1981
Earlier work this paper cites.
Robust decentralized regulation: a linear programming approach
J Bernussou, PLD Peres, and JC Geromel · 1989
Earlier work this paper cites.
State-space solutions to standard h 2 h_{2} and h ∞ h_{\infty} control problems
J.C. Doyle, K. Glover, P.P. Khargonekar, and B.A. Francis · 1989
Earlier work this paper cites.
Rates of convergence for empirical processes of stationary mixing sequences
Bin Yu · 1994
Earlier work this paper cites.
Mixed H2 /H ∞ \infty control
Carsten Scherer · 1995
Earlier work this paper cites.
Robust and Optimal Control , volume 40
Kemin Zhou, John Comstock Doyle, Keith Glover, et al · 1996
Earlier work this paper cites.
Multiobjective output-feedback control via LMI optimization
C. Scherer, P. Gahinet, and M. Chilali · 1997
Earlier work this paper cites.
A note on hurwitz stability of matrices
Guang-Ren Duan and Ron J Patton · 1998
Earlier work this paper cites.
LMI-based controller synthesis: A unified formulation and solution
Izumi Masubuchi, Atsumi Ohara, and Nobuhide Suda · 1998
Earlier work this paper cites.
Numerical optimization
Stephen Wright, Jorge Nocedal, et al · 1999
Earlier work this paper cites.
Extended kalman filtering and weighted least squares dynamic identification of robot
Maxime Gautier and Ph Poignet · 2001
Earlier work this paper cites.
Convex Optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Cited alongside, same era.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Cited alongside, same era.
Ordinary differential equations , page 1–2
Arieh Iserles · 2008
Cited alongside, same era.
Parameter estimation and model selection in computational biology
Gabriele Lillacci and Mustafa Khammash · 2010
Cited alongside, same era.
Cauchy and the gradient method
Claude Lemaréchal · 2012
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2014
Cited alongside, same era.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I. Jordan · 2018
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
LQR through the lens of first order methods: Discrete-time case
Jingjing Bu, Afshin Mesbahi, Maryam Fazel, and Mehran Mesbahi · 2019
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter Bartlett, and Martin Wainwright · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Józefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, Jonas Schneider, Szymon Sidor, Josh Tobin, Peter Welinder, Lilian Weng, and Wojciech Zaremba · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A Tropp · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Later among the works it cites.
Learning the globally optimal distributed LQ regulator
Luca Furieri, Yang Zheng, and Maryam Kamgarpour · 2020
Later among the works it cites.
Robustness analysis of non-convex stochastic gradient descent using biased expectations
Kevin Scaman and Cedric Malherbe · 2020
Later among the works it cites.
Policy optimization for ℋ 2 \mathcal{H}_{2} linear control with ℋ ∞ \mathcal{H}_{\infty} robustness guarantee: Implicit regularization and global convergence
Kaiqing Zhang, Bin Hu, and Tamer Basar · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Later among the works it cites.
Optimizing static linear feedback: Gradient method
Ilyas Fatkhullin and Boris Polyak · 2021
Later among the works it cites.
Distributed reinforcement learning for decentralized linear quadratic control: A derivative-free policy optimization approach
Yingying Li, Yujie Tang, Runyu Zhang, and Na Li · 2021
Later among the works it cites.
Convergence and sample complexity of gradient methods for the model-free linear quadratic regulator problem
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanovic · 2021
Later among the works it cites.
Learning optimal controllers by policy gradient: Global optimality via convex parameterization
Yue Sun and Maryam Fazel · 2021
Later among the works it cites.
Analysis of the optimization landscape of linear quadratic Gaussian (LQG) control
Yujie Tang, Yang Zheng, and Na Li · 2021
Later among the works it cites.