Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning attempts to find an optimal control action for an unknown dynamical system by directly searching over the parameter space of controllers.
1907
Earlier work this paper cites.
D. Kleinman, “On an iterative technique for Riccati equation computations,” IEEE Trans. Automat. Control , vol. 13, no. 1, pp. 114–115, 1968
1968
Earlier work this paper cites.
W. S. Levine and M. Athans, “On the determination of the optimal constant output feedback gains for linear multivariable systems,” IEEE Trans. Automat. Control , vol. 15, no. 1, pp. 44–48, 1970
1970
Earlier work this paper cites.
J. Ackermann, “Parameter space design of robust control systems,” IEEE Trans. Automat. Control , vol. 25, no. 6, pp. 1058–1072, 1980
1980
Earlier work this paper cites.
H. T. Toivonen, “A globally convergent algorithm for the optimal constant output feedback problem,” Int. J. Control , vol. 41, no. 6, pp. 1589–1599, 1985
1985
Earlier work this paper cites.
A. Vannelli and M. Vidyasagar, “Maximal Lyapunov functions and domains of attraction for autonomous nonlinear systems,” Automatica , vol. 21, no. 1, pp. 69 – 80, 1985
1985
Earlier work this paper cites.
H. T. Toivonen and P. M. Mäkilä, “Newton’s method for solving parametric linear quadratic control problems,” Int. J. Control , vol. 46, no. 3, pp. 897–911, 1987
1987
Earlier work this paper cites.
B. Anderson and J. Moore, Optimal Control; Linear Quadratic Methods . New York, NY: Prentice Hall, 1990
1990
Earlier work this paper cites.
E. Feron, V. Balakrishnan, S. Boyd, and L. El Ghaoui, “Numerical methods for H 2 H_{2} related problems,” in Proceedings of the 1992 American Control Conference , 1992, pp. 2921–2922
1992
Earlier work this paper cites.
P. L. D. Peres and J. C. Geromel, “An alternate numerical solution to the linear quadratic problem,” IEEE Trans. Automat. Control , vol. 39, no. 1, pp. 198–202, 1994
1994
Earlier work this paper cites.
H. K. Khalil, Nonlinear Systems . New York: Prentice Hall, 1996
1996
Earlier work this paper cites.
A. W. Vaart and J. A. Wellner, Weak convergence and empirical processes: with applications to statistics . Springer, 1996
1996
Earlier work this paper cites.
T. Rautert and E. W. Sachs, “Computational design of optimal output feedback controllers,” SIAM J. Optim , vol. 7, no. 3, pp. 837–852, 1997
1997
Earlier work this paper cites.
S.-I. Amari, “Natural gradient works efficiently in learning,” Neural Comput. , vol. 10, no. 2, pp. 251–276, 1998
1998
Earlier work this paper cites.
G. E. Dullerud and F. Paganini, A course in robust control theory: a convex approach . New York: Springer-Verlag, 2000
2000
Earlier work this paper cites.
V. Balakrishnan and L. Vandenberghe, “Semidefinite programming duality and linear time-invariant systems,” IEEE Trans. Automat. Control , vol. 48, no. 1, pp. 30–41, 2003
2003
Earlier work this paper cites.
2004
Cited alongside, same era.
S. Boyd and L. Vandenberghe, Convex optimization . Cambridge University Press, 2004
2004
Cited alongside, same era.
D. Bertsekas, “Approximate policy iteration: A survey and some new methods,” J. Control Theory Appl. , vol. 9, no. 3, pp. 310–335, 2011
2011
Cited alongside, same era.
F. Lin, M. Fardad, and M. R. Jovanović, “Augmented Lagrangian approach to design of structured optimal state feedback gains,” IEEE Trans. Automat. Control , vol. 56, no. 12, pp. 2923–2929, 2011
2011
Cited alongside, same era.
S. Bittanti, A. J. Laub, and J. C. Willems, The Riccati Equation . Berlin, Germany: Springer-Verlag, 2012
2012
Cited alongside, same era.
A. Nagabandi, G. Kahn, R. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in IEEE Int Conf. Robot. Autom. , 2018, pp. 7559–7566
2018
Later among the works it cites.
M. Simchowitz, H. Mania, S. Tu, M. I. Jordan, and B. Recht, “Learning without mixing: Towards a sharp analysis of linear system identification,” in Proc. Mach. Learn. Res. , 2018, pp. 439––473
2018
Later among the works it cites.
H. Mania, A. Guy, and B. Recht, “Simple random search of static linear policies is competitive for reinforcement learning,” in NeurIPS , vol. 31, 2018
2018
Later among the works it cites.
M. Fazel, R. Ge, S. M. Kakade, and M. Mesbahi, “Global convergence of policy gradient methods for the linear quadratic regulator,” in Proc. Int’l Conf. Machine Learning , 2018, pp. 1467–1476
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
F. Lin, M. Fardad, and M. R. Jovanović, “Design of optimal sparse feedback gains via the alternating direction method of multipliers,” IEEE Trans. Automat. Control , vol. 58, no. 9, pp. 2426–2431, 2013
2013
Cited alongside, same era.
B. Polyak, M. Khlebnikov, and P. Shcherbakov, “An LMI approach to structured sparse feedback design in linear control systems,” in Proceedings of the 2013 European Control Conference , 2013, pp. 833–838
2013
Cited alongside, same era.
M. Rudelson and R. Vershynin, “Hanson-Wright inequality and sub-Gaussian concentration,” Electron. Commun. Probab. , vol. 18, 2013
2013
Cited alongside, same era.
M. Ledoux and M. Talagrand, Probability in Banach Spaces: isoperimetry and processes . Springer Science & Business Media, 2013
2013
Cited alongside, same era.
N. K. Dhingra, M. R. Jovanović, and Z. Q. Luo, “An ADMM algorithm for optimal sensor and actuator selection,” in Proceedings of the 53rd IEEE Conference on Decision and Control , 2014, pp. 4039–4044
2014
Cited alongside, same era.
D. Pollard, “Mini empirical,” 2015. [Online]. Available: http://www.stat.yale.edu/~pollard/Books/Mini/
2015
Cited alongside, same era.
R. Vershynin, High-dimensional probability: An introduction with applications in data science . Cambridge University Press, 2018
2018
Later among the works it cites.
Y. Abbasi-Yadkori, N. Lazic, and C. Szepesvári, “Model-free linear quadratic control via reduction to expert prediction,” in Proc. Mach. Learn. Res. , vol. 89, 2019, pp. 3108–3117
2019
Closest in time.
D. Malik, A. Panajady, K. Bhatia, K. Khamaru, P. L. Bartlett, and M. J. Wainwright, “Derivative-free methods for policy optimization: Guarantees for linear-quadratic systems,” J. Mach. Learn. Res. , vol. 51, p. 1–51, 2019
2019
Closest in time.
B. Recht, “A tour of reinforcement learning: The view from continuous control,” Annu. Rev. Control Robot. Auton. Syst. , vol. 2, pp. 253–279, 2019
2019
Closest in time.
M. Soltanolkotabi, A. Javanmard, and J. D. Lee, “Theoretical insights into the optimization landscape of over-parameterized shallow neural networks,” IEEE Trans. Inf. Theory , vol. 65, no. 2, pp. 742–769, 2019
2019
Closest in time.
J. P. Jansch-Porto, B. Hu, and G. E. Dullerud, “Convergence guarantees of policy optimization methods for Markovian jump linear systems,” in Proceedings of the American Control Conference , 2020
2020
Closest in time.
K. Zhang, B. Hu, and T. Başar, “Policy optimization for ℋ 2 \mathcal{H}_{2} linear control with ℋ ∞ \mathcal{H}_{\infty} robustness guarantee: Implicit regularization and global convergence,” in Learning for Dynamics and Control , vol. 120, 2020
2020
Closest in time.
L. Furieri, Y. Zheng, and M. Kamgarpour, “Learning the globally optimal distributed LQ regulator,” in Learning for Dynamics and Control , 2020, pp. 287–297
2020
Closest in time.
A. Zare, H. Mohammadi, N. K. Dhingra, T. T. Georgiou, and M. R. Jovanović, “Proximal algorithms for large-scale statistical modeling and sensor/actuator selection,” IEEE Trans. Automat. Control , vol. 65, no. 8, pp. 3441–3456, 2020
2020
Closest in time.
H. Mohammadi, M. Soltanolkotabi, and M. R. Jovanović, “Random search for learning the linear quadratic regulator,” in Proceedings of the 2020 American Control Conference , 2020, pp. 4798–4803
2020
Closest in time.
H. Mohammadi, M. Soltanolkotabi, and M. R. Jovanović, “On the linear convergence of random search for discrete-time LQR,” IEEE Control Syst. Lett. , vol. 5, no. 3, pp. 989–994, July 2021
2021
Closest in time.
M. Fardad, F. Lin, and M. R. Jovanović, “Sparsity-promoting optimal control for a class of distributed systems,” in Proceedings of the 2011 American Control Conference , 2011, pp. 2050–2055
2055
Closest in time.