Fetching the paper…
Reading the bibliography…
Owing to the growth of interest in Reinforcement Learning in the last few years, gradient based policy control methods have been gaining popularity for Control problems as well.
M. Fragoso, “Discrete-time jump LQG problem,” International Journal of Systems Science, vol. 20, no. 12, pp. 2539–2545, 1989
1989
Earlier work this paper cites.
K. J. Astrom and B. Wittenmark. Adaptive Control. Addison Wesley, 1989
1989
Earlier work this paper cites.
Y. Bar-Shalom and X. Li, “Estimation and tracking- principles, tech- niques, and software,” Norwood, MA: Artech House, Inc, 1993., 1993
1993
Earlier work this paper cites.
Claude-Nicolas Fiechter. PAC adaptive control of liner systems. In Proceeding COLT ’94 Proceedings of the seventh annual conference on Computational learning theory, pages 88–97, 1994
1994
Earlier work this paper cites.
S.J. Bradtke, B.E. Ydstie, and a.G. Barto. Adaptive linear quadratic control using policy iteration. Proceedings of American Control Conference, 3(2):3475–3479, 1994. doi: 10.1109/ACC.1994.735224
1994
Earlier work this paper cites.
D. Sworder and J. Boyd, Estimation problems in hybrid systems. Cambridge University Press, 1999
1999
Earlier work this paper cites.
Lennart Ljung, editor. System Identification (2Nd Ed.): Theory for the User. Prentice Hall PTR, Upper Saddle River, NJ, USA, 1999. ISBN 0-13-656695-2
1999
Earlier work this paper cites.
V. Pavlovic, J. Rehg, and J. MacCormick, “Learning switching linear models of human motion,” in Advances in Neural Information Pro- cessing Systems, 2000. [Just cite]
2000
Earlier work this paper cites.
S. Kakade. A natural policy gradient. In NIPS, 2001
2001
Earlier work this paper cites.
Abraham D Flaxman, Adam Tauman Kalai, and H Brendan McMahan. Online convex optimization in the bandit setting: gradient descent without a gradient. In Proceedings of the sixteenth annual ACM-SIAM symposium on Discrete algorithms, pages 385–394. Society for Industrial and Applied Mathematics, 2005
2005
Cited alongside, same era.
O. Costa, M. Fragoso, and R. Marques, Discrete-time Markov jump linear systems. Springer London, 2006
2006
Cited alongside, same era.
M. Salathe, M. Kazandjieva, J. W. Lee, P. Levis, M. W. Feldman, and J. H. Jones, “A high resolution human contact network for infectious disease transmission,” Proceedings of the National Academy of Sciences, vol. 107, no. 51, pp. 22 020–22 025, 2010
2010
Cited alongside, same era.
2011
Cited alongside, same era.
K. Gopalakrishnan, H. Balakrishnan, and R. Jordan, “Stability of networked systems with switching topologies,” in IEEE Conference on Decision and Control, 2016, pp. 1889–1897
2016
Later among the works it cites.
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Do- minik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering the game of go with deep neural networks and tree search. Nature, 529, 2016
2016
Later among the works it cites.
B. Hu, P. Seiler, and A. Rantzer, “A unified analysis of stochastic optimization methods using jump system theory and quadratic con- straints,” in Conference on Learning Theory, 2017, pp. 1157–1189
2017
Later among the works it cites.
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu. On the sample complexity of the linear quadratic regulator. ArXiv e-prints, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Fox, E. S. M. Jordan, and A. Willsky, “Bayesian nonparametric inference of switching dynamic linear models,” IEEE Transactions on Signal Processing, vol. 59, no. 4, pp. 1569 – 1585, 2011
2011
Cited alongside, same era.
Yasin Abbasi-Yadkori and Csaba Szepesvari. Regret bounds for the adaptive control of linear quadratic systems. Conference on Learning Theory, 2011. ISSN 15337928
2011
Cited alongside, same era.
Dimitri P. Bertsekas. Approximate policy iteration: A survey and some new methods. Journal of Control Theory and Applications, 9(3):310–335, 2011. ISSN 16726340. doi: 10.1007/s11768-011-1005-3
2011
Cited alongside, same era.
A. N. Vargas, E. F. Costa, and J. B. R. do Val, “On the control of markov jump linear systems with no mode observation: application to a dc motor device,” Int. J. Robust. Nonlinear Control, vol. 23, no. 10, pp. 1136–1150, 2013
2013
Cited alongside, same era.
M. Zanin and F. Lillo, “Modelling the air transport with complex networks: A short review,” The European Physical Journal Special Topics, vol. 215, pp. 5–21, 2013
2013
Cited alongside, same era.
Cited in the paper.
2017
Later among the works it cites.
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht. Learning without mixing: Towards a sharp analysis of linear system identification. In COLT, 2018
2018
Later among the works it cites.
Sanjeev Arora, Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang. Towards provable control for unknown linear dynamical systems. 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Fazel, R. Ge, S. Kakade, and M. Mesbahi, “Global convergence of policy gradient methods for the linear quadratic regulator,” in Pro- ceedings of the 35th International Conference on Machine Learning, vol. 80, 2018, pp. 1467–1476
2018
Later among the works it cites.
B. Hu and U. Syed, “Characterizing the exact behaviors of temporal difference learning algorithms using markov jump linear system the- ory,” in Advances in Neural Information Processing Systems, 2019, pp. 8477–8488
2019
Later among the works it cites.