Fetching the paper…
Reading the bibliography…
We study the problem of controlling linear time-invariant systems with known noisy dynamics and adversarially chosen quadratic losses.
Linear optimal control
Anderson, B., Moore, J., and Molinari, B · 1972
Earlier work this paper cites.
On self tuning regulators
Åström, K. J. and Wittenmark, B · 1973
Earlier work this paper cites.
Reinforcement learning applied to linear quadratic regulation
Bradtke, S. J · 1993
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P · 1995
Earlier work this paper cites.
Robust and optimal control , volume 40
Zhou, K., Doyle, J. C., Glover, K., et al · 1996
Earlier work this paper cites.
Adaptive linear quadratic gaussian control: the cost-biased approach revisited
Campi, M. C. and Kumar, P · 1998
Earlier work this paper cites.
Semidefinite programming duality and linear time-invariant systems
Balakrishnan, V. and Vandenberghe, L · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Efficient algorithms for online decision problems
Kalai, A. and Vempala, S · 2005
Earlier work this paper cites.
Adaptive control of linear time invariant systems: the bet on the best principle
Bittanti, S. and Campi, M. C · 2006
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Abbeel, P., Coates, A., Quigley, M., and Ng, A. Y · 2007
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Auer, P. and Ortner, R · 2007
Earlier work this paper cites.
Constrained infinite-horizon linear quadratic regulation of discrete-time systems
Lee, J.-W. and Khargonekar, P. P · 2007
Cited alongside, same era.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2009
Cited alongside, same era.
Reinforcement learning and adaptive dynamic programming for feedback control
Lewis, F. L. and Vrabie, D · 2009
Cited alongside, same era.
Efficient computation of optimal actions
Todorov, E · 2009
Cited alongside, same era.
Markov decision processes with arbitrary reward processes
Yu, J. Y., Mannor, S., and Shimkin, N · 2009
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Cited alongside, same era.
Machine learning applications for data center optimization
Gao, J. and Jamidar, R · 2014
Later among the works it cites.
Linear controller design for chance constrained systems
Schildbach, G., Goulart, P., and Morari, M · 2015
Later among the works it cites.
Introduction to online convex optimization
Hazan, E · 2016
Later among the works it cites.
A semidefinite programming formulation of the lqr problem and its dual
Lee, D.-H. and Hu, J · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Later among the works it cites.
Thompson sampling for linear-quadratic control problems
Abeille, M. and Lazaric, A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret bounds for the adaptive control of linear quadratic systems
Abbasi-Yadkori, Y. and Szepesvári, C · 2011
Cited alongside, same era.
Efficient reinforcement learning for high dimensional linear quadratic systems
Ibrahimi, M., Javanmard, A., and Roy, B. V · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shalev-Shwartz, S · 2012
Cited alongside, same era.
Online learning in markov decision processes with adversarially chosen transition probability distributions
Abbasi, Y., Bartlett, P. L., Kanade, V., Seldin, Y., and Szepesvári, C · 2013
Cited alongside, same era.
Convex control design via covariance minimization
Dvijotham, K., Todorov, E., and Fazel, M · 2013
Cited alongside, same era.
Tracking adversarial targets
Abbasi-Yadkori, Y., Bartlett, P., and Kanade, V · 2014
Cited alongside, same era.
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S · 2017
Later among the works it cites.
Learning linear dynamical systems via spectral filtering
Hazan, E., Singh, K., and Zhang, C · 2017
Later among the works it cites.
Fast rates for online learning in linearly solvable markov decision processes
Neu, G. and Gómez, V · 2017
Later among the works it cites.
Robust policy search with applications to safe vehicle navigation
Sheckells, M., Garimella, G., and Kobilarov, M · 2017
Later among the works it cites.
Towards provable control for unknown linear dynamical systems
Arora, S., Hazan, E., Lee, H., Singh, K., Zhang, C., and Zhang, Y · 2018
Closest in time.
Global convergence of policy gradient methods for linearized control problems
Fazel, M., Ge, R., Kakade, S. M., and Mesbahi, M · 2018
Closest in time.