Fetching the paper…
Reading the bibliography…
We study the sample complexity of approximate policy iteration (PI) for the Linear Quadratic Regulator (LQR), building on a recent line of work using LQR as a testbed to understand the limits of reinforcement learning (RL) algorithms on continuous control tasks.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Incremental Dynamic Programming for On-Line Adaptive Optimal Control
Steven J. Bradtke · 1994
Earlier work this paper cites.
PAC Adaptive Control of Linear Systems
Claude-Nicolas Fiechter · 1997
Earlier work this paper cites.
Least-Squares Temporal Difference Learning
Justin Boyan · 1999
Earlier work this paper cites.
Linear Estimation
Thomas Kailath, Ali H. Sayed, and Babak Hassibi · 2000
Earlier work this paper cites.
Least-Squares Policy Iteration
Michail G. Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
On the Kronecker Product
Kathrin Schäcke · 2004
Earlier work this paper cites.
Relaxed Dynamic Programming
Bo Lincoln and Anders Rantzer · 2006
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Vol. II
Dimitri P. Bertsekas · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Invariant metrics, contractions and nonlinear matrix equations
Hosoo Lee and Yongdo Lim · 2008
Earlier work this paper cites.
An Analysis of Reinforcement Learning with Function Approximation
Francisco S. Melo, Sean P. Meyn, and M. Isabel Ribeiro · 2008
Earlier work this paper cites.
Regret Bounds for the Adaptive Control of Linear Quadratic Systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Online Least Squares Estimation with Self-Normalized Processes: An Application to Bandit Problems
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Efficient Reinforcement Learning for High Dimensional Linear Quadratic Systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin Van Roy · 2012
Cited alongside, same era.
Finite-Sample Analysis of Least-Squares Policy Iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Cited alongside, same era.
Hanson-Wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Cited alongside, same era.
Gaussian Measures
Vladimir I. Bogachev · 2015
Cited alongside, same era.
Regularized Policy Iteration with Nonparametric Function Spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Regret Bounds for Robust Adaptive Control of the Linear Quadratic Regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Later among the works it cites.
Finite Time Identification in Unstable Linear Systems
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis · 2018
Later among the works it cites.
Global Convergence of Policy Gradient Methods for the Linear Quadratic Regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Later among the works it cites.
Learning Without Mixing: Towards A Sharp Analysis of Linear System Identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I. Jordan, and Benjamin Recht · 2018
Later among the works it cites.
Least-Squares Temporal Difference Learning for the Linear Quadratic Regulator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Value and Policy Iterations in Optimal Control and Adaptive Dynamic Programming
Dimitri P. Bertsekas · 2017
Cited alongside, same era.
On the Sample Complexity of the Linear Quadratic Regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2017
Cited alongside, same era.
Random Gradient-Free Minimization of Convex Functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Cited alongside, same era.
Control of unknown linear systems with Thompson sampling
Yi Ouyang, Mukul Gagrani, and Rahul Jain · 2017
Cited alongside, same era.
Improved Regret Bounds for Thompson Sampling in Linear Quadratic Control Problems
Marc Abeille and Alessandro Lazaric · 2018
Cited alongside, same era.
Sharp bounds for the Lambert
Faris Alzahrani and Ahmed Salem · 2018
Cited alongside, same era.
Stephen Tu and Benjamin Recht · 2018
Later among the works it cites.
Model-Free Linear Quadratic Control via Reduction to Expert Prediction
Yasin Abbasi-Yadkori, Nevena Lazić, and Csaba Szepesvári · 2019
Closest in time.
Learning Linear-Quadratic Regulators Efficiently with only
Alon Cohen, Tomer Koren, and Yishay Mansour · 2019
Closest in time.
Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems
Dhruv Malik, Kush Bhatia, Koulik Khamaru, Peter L. Bartlett, , and Martin J. Wainwright · 2019
Closest in time.
Certainty Equivalent Control of LQR is Efficient
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Closest in time.
Near optimal finite time identification of arbitrary linear dynamical systems
Tuhin Sarkar and Alexander Rakhlin · 2019
Closest in time.
The Gap Between Model-Based and Model-Free Methods on the Linear Quadratic Regulator: An Asymptotic Viewpoint
Stephen Tu and Benjamin Recht · 2019
Closest in time.
Finite-Sample Analysis for SARSA and Q-Learning with Linear Function Approximation
Shaofeng Zou, Tengyu Xu, and Yingbin Liang · 2019
Closest in time.