Fetching the paper…
Reading the bibliography…
This manuscript surveys reinforcement learning from the perspective of optimization and control with a focus on continuous control applications.
About convergence of random search method in extremal control of multi-parameter systems
L. A. Rastrigin · 1963
Earlier work this paper cites.
When is a linear control system optimal?
R. E. Kalman · 1964
Earlier work this paper cites.
Optimal control of Markov processes with incomplete state information
K. J. Åström · 1965
Earlier work this paper cites.
Evolutionsstrategie und numerische Optimierung
H.-P. Schwefel · 1975
Earlier work this paper cites.
Bootstrap Methods: Another Look at the Jackknife
B. Efron · 1979
Earlier work this paper cites.
Discrete time stochastic adaptive control
G. C. Goodwin, P. J. Ramadge, and P. E. Caines · 1981
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovski and D. Yudin · 1983
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
R. S. Sutton · 1984
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
The complexity of Markov Decision Processes
C. H. Papadimitriou and J. N. Tsitsiklis · 1987
Earlier work this paper cites.
Learning to predict by the method of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Toward a theory of reinforcement-learning connectionist systems
R. J. Williams · 1988
Earlier work this paper cites.
Cooperative n-person Stackelberg games
W. F. Bialas · 1989
Earlier work this paper cites.
Reinforcement comparison
P. Dayan · 1991
Earlier work this paper cites.
The convergence of TD (
P. Dayan · 1992
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
J. C. Spall · 1992
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Adaptive linear quadratic control using policy iteration
S. J. Bradtke, B. E. Ydstie, and A. G. Barto · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
The Jackknife and Bootstrap
J. Shao and D. Tu · 1995
Earlier work this paper cites.
TD-gammon: A self-teaching backgammon program
G. Tesauro · 1995
Earlier work this paper cites.
Robust and Optimal Control
K. Zhou, J. C. Doyle, and K. Glover · 1995
Earlier work this paper cites.
Temporal differences-based policy iteration and applications in neuro-dynamic programming
D. P. Bertsekas and S. Ioffe · 1996
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
System Identification. Theory for the user
L. Ljung · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
A survey of computational complexity results in systems and control
V. D. Blondel and J. N. Tsitsiklis · 2000
Cited alongside, same era.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
Evolution Strategies - a comprehensive introduction
H.-G. Beyer and H.-P. Schwefel · 2002
Cited alongside, same era.
Finite Sample Properties of System Identification Methods
M. C. Campi and E. Weyer · 2002
Cited alongside, same era.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
C. Dann and E. Brunskill · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Later among the works it cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic policy gradient reinforcement learning on a simple 3D biped
R. Tedrake, T. W. Zhang, and H. S. Seung · 2004
Cited alongside, same era.
Online convex optimization in the bandit setting: gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2005
Cited alongside, same era.
Semi-supervised learning literature survey
X. Zhu · 2005
Cited alongside, same era.
A learning theory approach to system identification and stochastic adaptive control
M. Vidyasagar and R. L. Karandikar · 2008
Cited alongside, same era.
Convergence results for some temporal difference methods based on least squares
H. Yu and D. P. Bertsekas · 2009
Cited alongside, same era.
Optimal algorithms for online convex optimization with multi-point bandit feedback
A. Agarwal, O. Dekel, and L. Xiao · 2010
Cited alongside, same era.
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Later among the works it cites.
Planning for autonomous cars that leverage effects on human actions
D. Sadigh, S. Sastry, S. A. Seshia, and A. D. Dragan · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Later among the works it cites.
A system level approach to controller synthesis
Y.-S. Wang, N. Matni, and J. C. Doyle · 2016
Later among the works it cites.
Thompson Sampling for Linear-Quadratic Control Problems
M. Abeille and A. Lazaric · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause · 2017
Later among the works it cites.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 2017
Later among the works it cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2017
Later among the works it cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup · 2017
Later among the works it cites.
Game theoretic modeling of driver and vehicle interactions for verification and validation of autonomous vehicle control systems
N. Li, D. W. Oyler, M. Zhang, Y. Yildiz, I. Kolmanovsky, and A. R. Girard · 2017
Later among the works it cites.
Scalable system level synthesis for virtually localizable systems
N. Matni, Y.-S. Wang, and J. Anderson · 2017
Later among the works it cites.
A mathematical introduction to robotic manipulation
R. M. Murray · 2017
Later among the works it cites.
Random gradient-free minimization of convex functions
Y. Nesterov and V. Spokoiny · 2017
Later among the works it cites.
Learning-based control of unknown linear systems with thompson sampling
Y. Ouyang, M. Gagrani, and R. Jain · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
A. Rajeswaran, K. Lowrey, E. Todorov, and S. Kakade · 2017
Later among the works it cites.
Learning model predictive control for iterative tasks. a data-driven control framework
U. Rosolia and F. Borrelli · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, and I. Sutskever · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Y. Wu, E. Mansimov, S. Liao, R. Grosse, and J. Ba · 2017
Later among the works it cites.
Y. Abbasi-Yadkori, N. Lazic, and C. Szepesvári · 2018
Closest in time.
On the sample complexity of the linear quadratic regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2018
Closest in time.
Regret bounds for robust adaptive control of the linear quadratic regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
H. Mania, A. Guy, and B. Recht · 2018
Closest in time.
Learning without mixing
M. Simchowitz, H. Mania, S. Tu, M. I. Jordan, and B. Recht · 2018
Closest in time.
Least-squares temporal difference learning for the linear quadratic regulator
S. L. Tu and B. Recht · 2018
Closest in time.