Fetching the paper…
Reading the bibliography…
Model-free approaches for reinforcement learning (RL) and continuous control find policies based only on past states and rewards, without fitting a model of the system dynamics.
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems
T. L. Lai and C. Z. Wei · 1982
Earlier work this paper cites.
Theory and practice of recursive identification , volume 5
Lennart Ljung and Torsten Söderström · 1983
Earlier work this paper cites.
Optimal adaptive control and consistent parameter estimates for ARMAX model with quadratic cost
H. Chen and L. Guo · 1987
Earlier work this paper cites.
Asymptotically efficient self-tuning regulators
T. L. Lai and C. Z. Wei · 1987
Earlier work this paper cites.
Control oriented system identification: a worst-case/deterministic approach in H ∞ H_{\infty}
Arthur J Helmicki, Clas A Jacobson, and Carl N Nett · 1991
Earlier work this paper cites.
Adaptive linear quadratic control using policy iteration
Steven J Bradtke, B Erik Ydstie, and Andrew G Barto · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
PAC adaptive control of linear systems
C. Fiechter · 1997
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Adaptive linear quadratic gaussian control: the cost-biased approach revisited
M. C. Campi and P. R. Kumar · 1998
Earlier work this paper cites.
Average cost temporal-difference learning
John N. Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
Control-oriented system identification: an H ∞ H_{\infty} approach , volume 19
Jie Chen and Guoxiang Gu · 2000
Earlier work this paper cites.
Adaptive control of linear time invariant systems: The bet on the best principle
S. Bittanti and M.C. Campi · 2006
Earlier work this paper cites.
Prediction, learning, and games
Nicoló Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
A. L. Strehl, L. Li, E. Wiewiora, J. Langford, and M. L. Littman · 2006
Earlier work this paper cites.
Eigenvalue inequalities for matrix product
Fuzhen Zhang and Qingling Zhang · 2006
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Cited alongside, same era.
Rewarding Excursions: Extending Reinforcement Learning to Complex Domains
Istvan Szita · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
Forced-exploration based algorithms for playing in bandits with large action sets
Yasin Abbasi-Yadkori · 2009
Cited alongside, same era.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Cited alongside, same era.
Convergence results for some temporal difference methods based on least squares
Huizhen Yu and Dimitri P Bertsekas · 2009
Online decision-making with high-dimensional covariates
Hamsa Bastani and Mohsen Bayati · 2015
Later among the works it cites.
Time series analysis: forecasting and control
George EP Box, Gwilym M Jenkins, Gregory C Reinsel, and Greta M Ljung · 2015
Later among the works it cites.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Later among the works it cites.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Later among the works it cites.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2016
Later among the works it cites.
Introduction to online convex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Linearly parameterized bandits
Paat Rusmevichientong and John N. Tsitsiklis · 2010
Cited alongside, same era.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin V. Roy · 2012
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Cited alongside, same era.
Optimal control
Frank L Lewis, Draguna Vrabie, and Vassilis L Syrmos · 2012
Cited alongside, same era.
Regularized off-policy td-learning
Bo Liu, Sridhar Mahadevan, and Ji Liu · 2012
Cited alongside, same era.
Elad Hazan · 2016
Later among the works it cites.
Thompson sampling for linear-quadratic control problems
Marc Abeille and Alessandro Lazaric · 2017
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2017
Later among the works it cites.
Fast rates for online learning in linearly solvable markov decision processes
G. Neu and V. Gómez · 2017
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Daniel J. Russo, Zheng Wen, and Benjamin Van Roy · 2017
Later among the works it cites.
Learning-based control of unknown linear systems with Thompson sampling
Yi Ouyang, Mukul Gagrani, and Rahul Jain · 2017
Later among the works it cites.
Non-Asymptotic Analysis of Robust Control from Coarse-Grained Identification
S. Tu, R. Boczar, A. Packard, and B. Recht · 2017
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2017
Later among the works it cites.
Towards provable control for unknown linear dynamical systems
Sanjeev Arora, Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Closest in time.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Closest in time.
Global convergence of policy gradient methods for linearized control problems
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Closest in time.