Fetching the paper…
Reading the bibliography…
Machine and reinforcement learning (RL) are increasingly being applied to plan and control the behavior of autonomous systems interacting with the physical world.
R. E. Kalman, “Design of self-optimizing control system,” Trans. ASME , vol. 80, pp. 468–478, 1958
1958
Earlier work this paper cites.
R. E. Bellman, Adaptive control processes: a guided tour . Princeton university press, 2015, vol. 2045, original work published 1961
1961
Earlier work this paper cites.
K. J. Åström and P. Eykhoff, “System identification—a survey,” Automatica , vol. 7, no. 2, pp. 123–162, 1971
1971
Earlier work this paper cites.
H. Akaike, “Information theory and an extension of the maximum likelihood principle,” in 2nd Intl. Symposium on Information Theory, 1973 . Akademiai Kaido, 1973, pp. 267–281
1973
Earlier work this paper cites.
K. J. Åström and B. Wittenmark, “On self tuning regulators,” Automatica , vol. 9, no. 2, pp. 185–199, 1973
1973
Earlier work this paper cites.
G. C. Goodwin and R. L. Payne, Dynamic system identification: experiment design and data analysis . Academic press, 1977
1977
Earlier work this paper cites.
L. Ljung, “Analysis of recursive stochastic algorithms,” IEEE Trans. on Automatic Control , vol. AC-22, pp. 551–575, Jan. 1977
1977
Earlier work this paper cites.
B. Egardt, Stability of Adaptive controllers . Berlin, FRG: Springer-Verlag, Jan. 1979
1979
Earlier work this paper cites.
G. Stein, “Adaptive flight control: A pragmatic view,” in Applications of Adaptive Control . Elsevier, 1980, pp. 291–312
1980
Earlier work this paper cites.
K. J. Åström, “Why use adaptive techniques for steering large tankers?” Intl. Journ. of Control , vol. 32, no. 4, pp. 689–708, 1980
1980
Earlier work this paper cites.
G. Goodwin, P. Ramadge, and P. E. Caines, “Discrete time stochastic adaptive control,” SIAM Journ. Control & Opt. , vol. 19, no. 6, pp. 829–853, 1981
1981
Earlier work this paper cites.
T. Lai and H. Robbins, “Iterated least squares in multiperiod control,” Adv. in Applied Math. , vol. 3, no. 1, pp. 50–73, 1982
1982
Earlier work this paper cites.
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Neuronlike adaptive elements that can solve difficult learning control problems,” IEEE Trans. on systems, man, and cybernetics , no. 5, pp. 834–846, 1983
1983
Earlier work this paper cites.
P. A. Ioannou and P. V. Kokotovic, “Instability analysis and improvement of robustness of adaptive control,” Automatica , vol. 20, no. 5, pp. 583–594, 1984
1984
Earlier work this paper cites.
K. J. Åström and T. Hägglund, “Automatic tuning of simple regulators with specifications on phase and amplitude margins,” Automatica , vol. 20, no. 5, pp. 645–651, 1984
1984
Earlier work this paper cites.
C. Rohrs, L. Valavani, M. Athans et al. , “Robustness of continuous-time adaptive control algorithms in the presence of unmodeled dynamics,” IEEE Trans. on Automatic Control , vol. 30, no. 9, pp. 881–889, 1985
1985
Earlier work this paper cites.
T. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in Applied Mathematics , vol. 6, no. 1, pp. 4–2, 1985
1985
Earlier work this paper cites.
T. Lai, “Asymptotically efficient adaptive control in stochastic regression models,” Adv. in Applied Math. , vol. 7, no. 1, pp. 23 – 45, 1986
1986
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” ML , vol. 3, no. 1, pp. 9–44, 1988
1988
Earlier work this paper cites.
K. J. Åström and B. Wittenmark, Adaptive control . Courier Corporation, 2013, original work published 1989
1989
Earlier work this paper cites.
S. Sastry and M. Bodson, Adaptive control: stability, convergence and robustness . Courier Corporation, 2011, original work published 1989
1989
Earlier work this paper cites.
C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, King’s College, Cambridge, 1989
1989
Earlier work this paper cites.
H.-F. Chen and L. Guo, Identification and stochastic adaptive control . Springer Science & Business Media, 2012, originally published 1991
1991
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” ML , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” ML , vol. 8, no. 3-4, pp. 229–256, 1992
1992
Earlier work this paper cites.
E. Barnard, “Temporal-difference methods and Markov models,” IEEE Trans. on Systems, Man, and Cyber. , vol. 23, no. 2, pp. 357–365, 1993
1993
Earlier work this paper cites.
M. A. Dahleh, T. V. Theodosopoulos, and J. N. Tsitsiklis, “The sample complexity of worst-case identification of FIR linear systems,” Systems & Control Letters , vol. 20, no. 3, pp. 157–166, 1993
1993
Earlier work this paper cites.
G. A. Rummery and M. Niranjan, On-line Q-learning using connectionist systems . University of Cambridge, Department of Engineering Cambridge, England, 1994, vol. 37
1994
Earlier work this paper cites.
S. J. Bradtke, B. E. Ydstie, and A. G. Barto, “Adaptive linear quadratic control using policy iteration,” in 1994 American Control Conf. , vol. 3, June 1994, pp. 3475–3479 vol.3
1994
Earlier work this paper cites.
R. Braatz, P. Young, J. C. Doyle et al. , “Computational complexity of μ \mu calculation,” IEEE Trans. on Automatic Control , vol. 39, no. 5, pp. 1000–1002, 1994
1994
Earlier work this paper cites.
P. Mäkilä, J. R. Partington et al. , “Worst-case control-relevant identification,” Automatica , vol. 31, no. 12, pp. 1799–1819, 1995
1995
Earlier work this paper cites.
L. Guo, “Convergence and logarithm laws of self-tuning regulators,” Automatica , vol. 31, no. 3, pp. 435–450, 1995
1995
Earlier work this paper cites.
G. Tesauro, “Temporal difference learning and TD-gammon,” Comm.s of the ACM , vol. 38, no. 3, pp. 58–68, 1995
1995
Earlier work this paper cites.
A. G. Barto, S. J. Bradtke, and S. P. Singh, “Learning to act using real-time dynamic programming,” AI , vol. 72, no. 1-2, pp. 81–138, 1995
1995
Earlier work this paper cites.
V. Phansalkar and M. Thathachar, “Local and global opt. algorithms for generalized learning automata,” Neural Computation , vol. 7, no. 5, pp. 950–973, 1995
1995
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-dynamic programming . Athena Scientific Belmont, MA, 1996, vol. 5
1996
Earlier work this paper cites.
C.-N. Fiechter, “PAC adaptive control of linear systems,” in COLT , vol. 6, no. 09. Citeseer, 1997, pp. 72–80
1997
Earlier work this paper cites.
A. N. Burnetas and M. N. Katehakis, “Optimal adaptive policies for Markov decision processes,” Mathematics of Operations Research , vol. 22, no. 1, pp. 222–255, 1997
1997
Earlier work this paper cites.
T. L. Graves and T. L. Lai, “Asymptotically efficient adaptive choice of control laws in controlled Markov chains,” SIAM J. Control and Optimization , vol. 35, no. 3, pp. 715–743, 1997
1997
Earlier work this paper cites.
H. Hjalmarsson, M. Gevers, S. Gunnarsson et al. , “Iterative feedback tuning: theory and applications,” IEEE Control sys. mag. , vol. 18, no. 4, pp. 26–41, 1998
1998
Earlier work this paper cites.
L. Ljung, “System identification,” Wiley Encyclopedia of Electrical and Electronics Engineering , pp. 1–19, 1999
1999
Earlier work this paper cites.
M. C. Campi and E. Weyer, “Finite sample properties of system identification methods,” IEEE Trans. on Automatic Control , vol. 47, no. 8, pp. 1329–1334, 2002
2002
Cited alongside, same era.
M. Kearns and S. Singh, “Near-optimal reinforcement learning in polynomial time,” Machine Learning , vol. 49, no. 2-3, pp. 209–232, 2002
2002
Cited alongside, same era.
A. Nedić and D. P. Bertsekas, “Least squares policy evaluation algorithms with linear function approximation,” Discrete Event Dynamic Systems , vol. 13, no. 1-2, pp. 79–110, 2003
2003
Cited alongside, same era.
M. G. Lagoudakis and R. Parr, “Least-Squares Policy Iteration,” Journ. of ML Research , vol. 4, pp. 1107–1149, 2003
2003
Cited alongside, same era.
M. Vidyasagar and R. L. Karandikar, “A learning theory approach to system identification and stochastic adaptive control,” in Probabilistic and randomized methods for design under uncertainty . Springer, 2006, pp. 265–302
S. Levine, C. Finn, T. Darrell et al. , “End-to-end training of deep visuomotor policies,” JMLR , vol. 17, no. 1, pp. 1334–1373, Jan. 2016
2016
Later among the works it cites.
D. Silver, A. Huang, C. J. Maddison et al. , “Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, pp. 484–489, 01 2016
2016
Later among the works it cites.
E. Kaufmann, O. Cappé, and A. Garivier, “On the complexity of best-arm identification in multi-armed bandit models,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 1–42, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2006
Cited alongside, same era.
D. A. Bristow, M. Tharayil, and A. G. Alleyne, “A survey of iterative learning control,” IEEE ctrl sys. mag. , vol. 26, no. 3, pp. 96–114, 2006
2006
Cited alongside, same era.
A. L. Strehl, L. Li, E. Wiewiora et al. , “PAC model-free reinforcement learning,” in 23rd Intl. conference on ML . ACM, 2006, pp. 881–888
2006
Cited alongside, same era.
A. Astolfi, D. Karagiannis, and R. Ortega, Nonlinear and adaptive control with applications . Springer Science & Business Media, 2007
2007
Cited alongside, same era.
W. B. Powell, Approximate Dynamic Programming: Solving the curses of dimensionality . John Wiley & Sons, 2007, vol. 703
2007
Cited alongside, same era.
P. Auer and R. Ortner, “Logarithmic online regret bounds for undiscounted reinforcement learning,” in Advances in Neural Information Processing Systems 19 , 2007
2007
Cited alongside, same era.
J.-Y. Audibert, R. Munos, and C. Szepesvári, “Exploration-exploitation tradeoff using variance estimates in multi-armed bandits,” Theor. Comput. Sci. , vol. 410, no. 19, pp. 1876–1902, Apr. 2009. [Online]. Available: http://dx.doi.org/10.1016/j.tcs.2009.01.016
2009
Cited alongside, same era.
P. Auer, T. Jaksch, and R. Ortner, “Near-optimal regret bounds for reinforcement learning,” in Advances in Neural Information Processing Systems 22 , 2009
2009
Cited alongside, same era.
2017
Later among the works it cites.
Y. Ouyang, M. Gagrani, and R. Jain, “Control of unknown linear systems with thompson sampling,” in 2017 55th Allerton Conf. on Comm., Control, and Computing , Oct 2017, pp. 1198–1205
2017
Later among the works it cites.
C. Dann, T. Lattimore, and E. Brunskill, “Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning,” in NeurIPS , 2017, pp. 5713–5723
2017
Later among the works it cites.
S. Agrawal and R. Jia, “Posterior sampling for reinforcement learning: worst-case regret bounds,” in Advances in Neural Information Processing Systems 31 , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
N. Matni, Y.-S. Wang, and J. Anderson, “Scalable system level synthesis for virtually localizable systems,” in 2017 IEEE 56th Conf. on Decision and Control (CDC) . IEEE, 2017, pp. 3473–3480
2017
Later among the works it cites.
2017
Later among the works it cites.
D. Silver, T. Hubert, J. Schrittwieser et al. , “A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play,” Science , vol. 362, no. 6419, pp. 1140–1144, 2018
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Later among the works it cites.
H. Mania, A. Guy, and B. Recht, “Simple random search of static linear policies is competitive for reinforcement learning,” in NeurIPS , 2018, pp. 1800–1809
2018
Later among the works it cites.
M. Hardt, T. Ma, and B. Recht, “Gradient descent learns linear dynamical systems,” Journ. of ML Research , vol. 19, no. 1, pp. 1025–1068, 2018
2018
Later among the works it cites.
M. Simchowitz, H. Mania, S. Tu et al. , “Learning without mixing: Towards a sharp analysis of linear system identification,” in Conf. On Learning Theory , 2018, pp. 439–473
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Fattahi and S. Sojoudi, “Data-driven sparse system identification,” in 2018 56th Allerton Conf. on Comm., Control, and Computing . IEEE, 2018, pp. 462–469
2018
Later among the works it cites.
D. J. Russo, B. Van Roy, A. Kazerouni et al. , “A tutorial on thompson sampling,” Foundations and Trends on ML , vol. 11, no. 1, pp. 1–96, Jul. 2018
2018
Later among the works it cites.
M. Abeille and A. Lazaric, “Improved regret bounds for thompson sampling in linear quadratic control problems,” in Intl. Conf. on ML , 2018, pp. 1–9
2018
Later among the works it cites.
S. Dean, H. Mania, N. Matni et al. , “Regret bounds for robust adaptive control of the linear quadratic regulator,” in NeurIPS , 2018, pp. 4188–4197
2018
Later among the works it cites.
A. Rantzer, “Concentration bounds for single parameter adaptive control,” in 2018 American Control Conf. IEEE, 2018, pp. 1862–1866
2018
Later among the works it cites.
B. Recht, “A tour of reinforcement learning: The view from continuous control,” Review of Control, Robotics, and Autonomous Sys. , 2018
2018
Later among the works it cites.
S. Tu and B. Recht, “Least-squares temporal difference learning for the linear quadratic regulator,” in Intl. Conf. on ML , 2018, pp. 5012–5021
2018
Later among the works it cites.
M. Fazel, R. Ge, S. Kakade et al. , “Global convergence of policy gradient methods for the linear quadratic regulator,” in Proceedings of the 35th Intl. Conf. on ML , vol. 80. PMLR, 10–15 Jul 2018, pp. 1467–1476
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Garivier, P. Menard, and G. Stoltz, “Explore first, exploit next: The true shape of regret in bandit problems,” Mathematics of Operations Research , Jun. 2018
2018
Later among the works it cites.
J. Ok, A. Proutiere, and D. Tranos, “Exploration in structured reinforcement learning,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle et al. , Eds. Curran Associates, Inc., 2018, pp. 8874–8882. [Online]. Available: http://papers.nips.cc/paper/8103-exploration-in-structured-reinforcement-learning.pdf
2018
Later among the works it cites.
A. Sidford, M. Wang, X. Wu et al. , “Near-optimal time and sample complexities for solving Markov decision processes with a generative model,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle et al. , Eds. Curran Associates, Inc., 2018, pp. 5186–5196
2018
Later among the works it cites.
C. Jin, Z. Allen-Zhu, S. Bubeck et al. , “Is Q-learning provably efficient?” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle et al. , Eds. Curran Associates, Inc., 2018, pp. 4863–4873
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
Y. Abbasi-Yadkori, N. Lazic, and C. Szepesvári, “Model-free linear quadratic control via reduction to expert prediction,” in The 22nd Intl. Conf. on AI and Statistics , 2019, pp. 3108–3117
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
D. Malik, K. Bhatia, K. Khamaru et al. , “Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems,” in AISTATS , 2019
2019
Closest in time.
2019
Closest in time.