Fetching the paper…
Reading the bibliography…
Model-free learning-based control methods have seen great success recently.
Yingying Li, Yujie Tang, Runyu Zhang, and Na Li · 1912
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Applied nonlinear control , volume 199
Jean-Jacques E Slotine, Weiping Li, et al · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Adaptive linear quadratic control using policy iteration
Steven J Bradtke, B Erik Ydstie, and Andrew G Barto · 1994
Earlier work this paper cites.
Robust and optimal control , volume 40
Kemin Zhou, John Comstock Doyle, Keith Glover, et al · 1996
Earlier work this paper cites.
System identification: theory for the user
Ljung Lennart · 1999
Earlier work this paper cites.
System identification
Lennart Ljung · 1999
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 2005
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Dynamic programming and optimal control 3rd edition, volume ii
Dimitri P Bertsekas · 2011
Earlier work this paper cites.
Feedback control theory
John C Doyle, Bruce A Francis, and Allen R Tannenbaum · 2013
Earlier work this paper cites.
A course in robust control theory: a convex approach , volume 36
Geir E Dullerud and Fernando Paganini · 2013
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity, 2014
Sébastien Bubeck · 2014
Earlier work this paper cites.
Nonlinear Control Systems Design 1989: Selected Papers from the IFAC Symposium, Capri, Italy, 14-16 June 1989
Alberto Isidori · 2014
Earlier work this paper cites.
Robust control of uncertain systems: Classical results and recent developments
Ian R Petersen and Roberto Tempo · 2014
Earlier work this paper cites.
From Small Signal to Exact Linearization of Swing Equations , chapter 3, pages 57–86
Abdelkrim Benchaib · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies, 2015
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Cited alongside, same era.
Mbmf: Model-based priors for model-free reinforcement learning
Somil Bansal, Roberto Calandra, Kurtland Chua, Sergey Levine, and Claire Tomlin · 2017
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2017
Cited alongside, same era.
Stephen Tu and Benjamin Recht · 2018
Later among the works it cites.
Navid Azizan, Sahin Lale, and Babak Hassibi · 2019
Later among the works it cites.
LQR through the lens of first order methods: Discrete-time case
Jingjing Bu, Afshin Mesbahi, Maryam Fazel, and Mehran Mesbahi · 2019
Later among the works it cites.
Learning linear-quadratic regulators efficiently with only
Alon Cohen, Tomer Koren, and Yishay Mansour · 2019
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis · 2017
Cited alongside, same era.
Prediction and control with temporal segment models
Nikhil Mishra, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Cited alongside, same era.
Learning-based control of unknown linear systems with thompson sampling
Yi Ouyang, Mukul Gagrani, and Rahul Jain · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Cited alongside, same era.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning, 2017
Théophane Weber, Sébastien Racanière, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, Demis Hassabis, David Silver, and Daan Wierstra · 2017
Cited alongside, same era.
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2019
Later among the works it cites.
Residual reinforcement learning for robot control
Tobias Johannink, Shikhar Bahl, Ashvin Nair, Jianlan Luo, Avinash Kumar, Matthias Loskyll, Juan Aparicio Ojea, Eugen Solowjow, and Sergey Levine · 2019
Later among the works it cites.
Finite-time analysis of approximate policy iteration for the linear quadratic regulator
Karl Krauth, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Modeling of inverted pendulum system with gravitational search algorithm optimized controller
Mohamed Magdy, Abdallah El Marhomy, and Mahmoud A Attia · 2019
Later among the works it cites.
Certainty equivalent control of LQR is efficient
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2019
Later among the works it cites.
Non-asymptotic identification of lti systems from a single trajectory
Samet Oymak and Necmiye Ozay · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Analyzing the variance of policy gradient estimators for the linear-quadratic regulator, 2019
James A. Preiss, Sébastien M. R. Arnold, Chen-Yu Wei, and Marius Kloft · 2019
Later among the works it cites.
A tour of reinforcement learning: The view from continuous control
Benjamin Recht · 2019
Later among the works it cites.
Finite-time system identification for partially observed lti systems of unknown order
Tuhin Sarkar, Alexander Rakhlin, and Munther A Dahleh · 2019
Later among the works it cites.
Neural lander: Stable drone landing control using learned dynamics
Guanya Shi, Xichen Shi, Michael O’Connell, Rose Yu, Kamyar Azizzadenesheli, Animashree Anandkumar, Yisong Yue, and Soon-Jo Chung · 2019
Later among the works it cites.
Uncertainty-aware model-based policy optimization
Tung-Long Vuong and Kenneth Tran · 2019
Later among the works it cites.
Feedback linearization for unknown systems via reinforcement learning, 2019
Tyler Westenbroek, David Fridovich-Keil, Eric Mazumdar, Shreyas Arora, Valmik Prabhu, S. Shankar Sastry, and Claire J. Tomlin · 2019
Later among the works it cites.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Zhuoran Yang, Yongxin Chen, Mingyi Hong, and Zhaoran Wang · 2019
Later among the works it cites.
Naive exploration is optimal for online LQR
Max Simchowitz and Dylan J Foster · 2020
Closest in time.
Improper learning for non-stochastic control
Max Simchowitz, Karan Singh, and Elad Hazan · 2020
Closest in time.
Deep reinforcement learning-based robust protection in electric distribution grids
Dongqi Wu, Dileep Kalathil, and Le Xie · 2020
Closest in time.