Fetching the paper…
Reading the bibliography…
We consider the problem of learning in Linear Quadratic Control systems whose transition parameters are initially unknown.
A bound on tail probabilities for quadratic forms in independent random variables
David Lee Hanson and Farroll Tim Wright · 1971
Earlier work this paper cites.
Farrol Tim Wright · 1973
Earlier work this paper cites.
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems
Tze Leung Lai, Ching Zong Wei, et al · 1982
Earlier work this paper cites.
Optimal adaptive control of linear-quadratic-gaussian systems
PR Kumar · 1983
Earlier work this paper cites.
A survey of some results in stochastic adaptive control
Panqanamala Ramana Kumar · 1985
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, Tong Zhang, et al · 2012
Earlier work this paper cites.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin V Roy · 2012
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2017
Cited alongside, same era.
Optimism-based adaptive regulation of linear-quadratic systems
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis · 2017
Cited alongside, same era.
Learning linear dynamical systems via spectral filtering
Elad Hazan, Karan Singh, and Cyril Zhang · 2017
Cited alongside, same era.
Control of unknown linear systems with thompson sampling
Input perturbations for adaptive regulation and learning
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Spectral filtering for general linear dynamical systems
Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Later among the works it cites.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
Alon Cohen, Tomer Koren, and Yishay Mansour · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Ouyang, Mukul Gagrani, and Rahul Jain · 2017
Cited alongside, same era.
Improved regret bounds for thompson sampling in linear quadratic control problems
Marc Abeille and Alessandro Lazaric · 2018
Cited alongside, same era.
Online linear quadratic control
Alon Cohen, Avinatan Hasidim, Tomer Koren, Nevena Lazic, Yishay Mansour, and Kunal Talwar · 2018
Cited alongside, same era.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz
Cited in the paper.
Model-free linear quadratic control via reduction to expert prediction
Yasin Abbasi-Yadkori, Nevena Lazic, and Csaba Szepesvari
Cited in the paper.
Online control with adversarial disturbances
Naman Agarwal, Brian Bullins, Elad Hazan, Sham Kakade, and Karan Singh
Cited in the paper.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter Bartlett, and Martin Wainwright · 2019
Later among the works it cites.
Certainty equivalent control of lqr is efficient
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Near optimal finite time identification of arbitrary linear dynamical systems
Tuhin Sarkar and Alexander Rakhlin · 2019
Later among the works it cites.
Naive exploration is optimal for online lqr
Max Simchowitz and Dylan J Foster · 2020
Closest in time.