Fetching the paper…
Reading the bibliography…
We consider the problem of online adaptive control of the linear quadratic regulator, where the true system parameters are unknown.
Deux remarques sur l’estimation
Patrice Assouad · 1983
Earlier work this paper cites.
Existence and comparison theorems for algebraic riccati equations for continuous-and discrete-time systems
ACM Ran and R Vreugdenhil · 1988
Earlier work this paper cites.
Bin Yu · 1997
Earlier work this paper cites.
Singular values and eigenvalues of non-hermitian block toeplitz matrices
Paolo Tilli · 1998
Earlier work this paper cites.
Approximate planning in large pomdps via reusable trajectories
Michael J Kearns, Yishay Mansour, and Andrew Y Ng · 2000
Earlier work this paper cites.
Competitive on-line statistics
Volodya Vovk · 2001
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael J Kearns, and John Langford · 2003
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Vol. I
Dimitri P Bertsekas · 2005
Earlier work this paper cites.
Relaxing dynamic programming
Bo Lincoln and Anders Rantzer · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Lecture 13: Linear quadratic lyapunov theory
Stephen Boyd · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
On the fundamental limits of adaptive sensing
Ery Arias-Castro, Emmanuel J Candes, and Mark A Davenport · 2012
Cited alongside, same era.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, Tong Zhang, et al · 2012
Cited alongside, same era.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson, Roman Vershynin, et al · 2013
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Online nonparametric regression
Alexander Rakhlin and Karthik Sridharan · 2014
Cited alongside, same era.
A note on the hanson-wright inequality for random vectors with dependencies
Radoslaw Adamczak et al · 2015
Cited alongside, same era.
Improved regret bounds for thompson sampling in linear quadratic control problems
Marc Abeille and Alessandro Lazaric · 2018
Later among the works it cites.
Lyapunov theory for discrete time systems
Nicoletta Bof, Ruggero Carli, and Luca Schenato · 2018
Later among the works it cites.
Online linear quadratic control
Alon Cohen, Avinatan Hasidim, Tomer Koren, Nevena Lazic, Yishay Mansour, and Kunal Talwar · 2018
Later among the works it cites.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Thompson sampling for linear-quadratic control problems
Marc Abeille and Alessandro Lazaric · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2018
Later among the works it cites.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
Alon Cohen, Tomer Koren, and Yishay Mansour · 2019
Later among the works it cites.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Near optimal finite time identification of arbitrary linear dynamical systems
Tuhin Sarkar and Alexander Rakhlin · 2019
Later among the works it cites.
Finite-Time System Identification for Partially Observed LTI Systems of Unknown Order
Tuhin Sarkar, Alexander Rakhlin, and Munther A. Dahleh · 2019
Later among the works it cites.
Learning linear dynamical systems with semi-parametric least squares
Max Simchowitz, Ross Boczar, and Benjamin Recht · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Closest in time.