Fetching the paper…
Reading the bibliography…
We present the first computationally-efficient algorithm with $\widetilde O(\sqrt{T})$ regret for learning in Linear Quadratic Control systems with unknown dynamics.
Weighted sums of certain dependent random variables
Kazuoki Azuma · 1967
Earlier work this paper cites.
A bound on tail probabilities for quadratic forms in independent random variables
David Lee Hanson and Farroll Tim Wright · 1971
Earlier work this paper cites.
A bound on tail probabilities for quadratic forms in independent random variables whose distributions are not necessarily symmetric
Farrol Tim Wright · 1973
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
A state constrained optimal control problem related to the sterilization of canned foods
Alfredo Bermúdez and Aurea Martinez · 1994
Earlier work this paper cites.
Optimal control: basics and beyond
Peter Whittle · 1996
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas, Dimitri P Bertsekas, Dimitri P Bertsekas, and Dimitri P Bertsekas · 2005
Earlier work this paper cites.
Optimal control models in finance
Ping Chen and Sardar MN Islam · 2005
Earlier work this paper cites.
Optimal control with engineering applications
Hans P Geering · 2007
Earlier work this paper cites.
Optimal control applied to biological models
Suzanne Lenhart and John T Workman · 2007
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, Tong Zhang, et al · 2012
Cited alongside, same era.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin V Roy · 2012
Cited alongside, same era.
Learning-based control of unknown linear systems with thompson sampling
Yi Ouyang, Mukul Gagrani, and Rahul Jain · 2017
Later among the works it cites.
Regret bounds for model-free linear quadratic control
Yasin Abbasi-Yadkori, Nevena Lazic, and Csaba Szepesvari · 2018
Later among the works it cites.
Improved regret bounds for thompson sampling in linear quadratic control problems
Marc Abeille and Alessandro Lazaric · 2018
Later among the works it cites.
Towards provable control for unknown linear dynamical systems, 2018
Sanjeev Arora, Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Later among the works it cites.
Online linear quadratic control
Alon Cohen, Avinatan Hassidim, Tomer Koren, Nevena Lazic, Yishay Mansour, and Kunal Talwar · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thompson sampling for linear-quadratic control problems
Marc Abeille and Alessandro Lazaric · 2017
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2017
Cited alongside, same era.
Finite time analysis of optimal adaptive policies for linear-quadratic systems
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis · 2017
Cited alongside, same era.
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter L Bartlett, and Martin J Wainwright · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Later among the works it cites.