Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has been successfully used to solve many continuous control tasks.
Reinforcement Learning Applied to Linear Quadratic Regulation
S. J. Bradtke · 1993
Earlier work this paper cites.
Incremental Dynamic Programming for On-Line Adaptive Optimal Control
S. J. Bradtke · 1994
Earlier work this paper cites.
Rates of Convergence for Empirical Processes of Stationary Mixing Sequences
B. Yu · 1994
Earlier work this paper cites.
Robust and Optimal Control
K. Zhou, J. C. Doyle, and K. Glover · 1995
Earlier work this paper cites.
Linear Least-Squares Algorithms for Temporal Difference Learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
An Analysis of Temporal-Difference Learning with Function Approximation
J. N. Tsitsiklis and B. V. Roy · 1997
Earlier work this paper cites.
Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Least-Squares Temporal Difference Learning
J. Boyan · 1999
Earlier work this paper cites.
Nonasymptotic bounds for autoregressive time series modeling
A. Goldenshluger and A. Zeevi · 2001
Earlier work this paper cites.
Least-Squares Policy Iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Stochastic Policy Gradient Reinforcement Learning on a Simple 3D Biped
R. Tedrake, T. W. Zhang, and H. S. Seung · 2004
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Vol. II
D. P. Bertsekas · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Rademacher Complexity Bounds for Non-I.I.D. Processes
M. Mohri and A. Rostamizadeh · 2008
Earlier work this paper cites.
Smallest Singular Value of a Random Rectangular Matrix
M. Rudelson and R. Vershynin · 2009
Earlier work this paper cites.
Convergence Results for Some Temporal Difference Methods Based on Least Squares
H. Yu and D. P. Bertsekas · 2009
Earlier work this paper cites.
Stability Bounds for Stationary
M. Mohri and A. Rostamizadeh · 2010
Earlier work this paper cites.
Sharp bounds on the rate of convergence of empirical covariance matrix
R. Adamczak, A. E. Litvak, A. Pajor, and N. Tomczak-Jaegermann · 2011
Earlier work this paper cites.
Hanson-Wright inequality and sub-gaussian concentration
M. Rudelson and R. Vershynin · 2011
Cited alongside, same era.
Freedman’s inequality for matrix martingales
J. A. Tropp · 2011
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2011
Cited alongside, same era.
Finite-Sample Analysis of Least-Squares Policy Iteration
A. Lazaric, M. Ghavamzadeh, and R. Munos · 2012
Cited alongside, same era.
Regularized Off-Policy TD-Learning
B. Liu, S. Mahadevan, and J. Liu · 2012
Cited alongside, same era.
The Generalization Ability of Online Algorithms for Dependent Data
A. Agarwal and J. C. Duchi · 2013
Cited alongside, same era.
The MOSEK optimization toolbox for MATLAB manual. Version 7.1 (Revision 28)
MOSEK ApS · 2015
Later among the works it cites.
Trust Region Policy Optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Later among the works it cites.
CVXPY: A Python-embedded modeling language for convex optimization
S. Diamond and S. Boyd · 2016
Later among the works it cites.
Regularized Policy Iteration with Nonparametric Function Spaces
A. Farahmand, M. Ghavamzadeh, C. Szepesvári, and S. Mannor · 2016
Later among the works it cites.
Continuous Deep Q-Learning with Model-based Acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Later among the works it cites.
Time Series Prediction and Online Learning
V. Kuznetsov and M. Mohri · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bandits With Heavy Tail
S. Bubeck, N. Cesa-Bianchi, and G. Lugosi · 2013
Cited alongside, same era.
Reinforcement Learning in Robotics: A Survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Cited alongside, same era.
Bounding the smallest singular value of a random matrix without concentration
V. Koltchinskii and S. Mendelson · 2013
Cited alongside, same era.
Covariance estimation for distributions with
N. Srivastava and R. Vershynin · 2013
Cited alongside, same era.
Learning Complex Neural Network Policies with Trajectory Optimization
S. Levine and V. Koltun · 2014
Cited alongside, same era.
On the singular values of random matrices
S. Mendelson and G. Paouris · 2014
Cited alongside, same era.
End-to-End Training of Deep Visuomotor Policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Later among the works it cites.
S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Later among the works it cites.
Prioritized Experience Replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Later among the works it cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Later among the works it cites.
On the Sample Complexity of the Linear Quadratic Regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2017
Closest in time.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
S. Gu, T. Lillicrap, Z. Ghahramani, R. E. Turner, and S. Levine · 2017
Closest in time.
Occupy the Cloud: Distributed Computing for the 99%
E. Jonas, Q. Pu, S. Venkataraman, I. Stoica, and B. Recht · 2017
Closest in time.
DDCO: Discovery of Deep Continuous Options for Robot Learning from Demonstrations
S. Krishnan, R. Fox, I. Stoica, and K. Goldberg · 2017
Closest in time.
Generalization bounds for non-stationary mixing processes
V. Kuznetsov and M. Mohri · 2017
Closest in time.
Nonparametric Risk Bounds for Time-Series Forecasting
D. J. McDonald, C. R. Shalizi, and M. Schervish · 2017
Closest in time.