Fetching the paper…
Reading the bibliography…
Recent progress in reinforcement learning has led to remarkable performance in a range of applications, but its deployment in high-stakes settings remains quite rare.
Finite-time system identification for partially observed lti systems of unknown order
Sarkar, T., Rakhlin, A., and Dahleh, M. A. (2019) · 1902
Earlier work this paper cites.
Reinforcement learning in non-stationary environments
Padakandla, S., Bhatnagar, S., et al. (2019) · 1905
Earlier work this paper cites.
Grünwald, P., de Heide, R., and Koolen, W. (2019) · 1906
Earlier work this paper cites.
Stabilization controllability and observability of linear autonomous systems
Hautus, M. (1970) · 1970
Earlier work this paper cites.
On the discrete time algebraic riccati equation
Payne, H. and Silverman, L. (1973) · 1973
Earlier work this paper cites.
Iterated least squares in multiperiod control
Lai, T. and Robbins, H. (1982) · 1982
Earlier work this paper cites.
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems
Lai, T. L. and Wei, C. Z. (1982) · 1982
Earlier work this paper cites.
Adaptive control with the stochastic approximation algorithm: Geometry and convergence
Becker, A., Kumar, P., and Wei, C.-Z. (1985) · 1985
Earlier work this paper cites.
Prediction mean squared error for state space models with estimated parameters
Kohn, R. and Ansley, C. F. (1986) · 1986
Earlier work this paper cites.
Asymptotically efficient adaptive control in stochastic regression models
Lai, T. (1986) · 1986
Earlier work this paper cites.
Extended least squares and their applications to adaptive control and prediction in linear systems
Lai, T. and Wei, C.-Z. (1986) · 1986
Earlier work this paper cites.
The astrom-wittenmark self-tuning regulator revisited and els-based adaptive trackers
Guo, L. and Chen, H.-F. (1991) · 1991
Earlier work this paper cites.
Asymptotic distributions of regression and autoregression coefficients with martingale difference disturbances
Anderson, T. W. and Kunitomo, N. (1992) · 1992
Earlier work this paper cites.
Convergence and logarithm laws of self-tuning regulators
Guo, L. (1995) · 1995
Earlier work this paper cites.
System Identification: Theory for the User
Ljung, L. (1997) · 1997
Earlier work this paper cites.
Lqr design with eigenstructure assignment capability [and application to aircraft flight control]
Choi, J. W. and Seo, Y. B. (1999) · 1999
Earlier work this paper cites.
Sequential analysis: some classical problems and new challenges
Lai, T. L. (2001) · 2001
Earlier work this paper cites.
Naive exploration is optimal for online lqr
Simchowitz, M. and Foster, D. J. (2020) · 2001
Earlier work this paper cites.
Logarithmic regret for learning linear quadratic regulators efficiently
Cassel, A., Cohen, A., and Koren, T. (2020) · 2002
Earlier work this paper cites.
Non-asymptotic and accurate learning of nonlinear dynamical systems
Sattar, Y. and Oymak, S. (2020) · 2002
Earlier work this paper cites.
Online learning of the kalman filter with logarithmic regret
Tsiamis, A. and Pappas, G. (2020) · 2002
Earlier work this paper cites.
Logarithmic regret for adversarial online control
Foster, D. J. and Simchowitz, M. (2020) · 2003
Cited alongside, same era.
Nonlinear system identification with prior knowledge of the region of attraction
Khosravi, M. and Smith, R. S. (2020) · 2003
Cited alongside, same era.
Logarithmic regret bound in partially observable linear dynamical systems
Lale, S., Azizzadenesheli, K., Hassibi, B., and Anandkumar, A. (2020) · 2003
Cited alongside, same era.
Application of adaptive lqr with repetitive control for ups systems
Shabaani, K. and Jalili-Kharaajoo, M. (2003) · 2003
Cited alongside, same era.
The kernel recursive least-squares algorithm
Engel, Y., Mannor, S., and Meir, R. (2004) · 2004
Cited alongside, same era.
Optimal experiment design for identification of arx models with constrained output in non-gaussian noise
Stojanovic, V., Nedic, N., Prsic, D., and Dubonjic, L. (2016) · 2016
Later among the works it cites.
Thompson sampling for linear-quadratic control problems
Abeille, M. and Lazaric, A. (2017) · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A. (2017) · 2017
Later among the works it cites.
Finite time analysis of optimal adaptive policies for linear-quadratic systems
Faradonbeh, M. K. S., Tewari, A., and Michailidis, G. (2017) · 2017
Later among the works it cites.
Adaptive input design for lti systems
Gerencsér, L., Hjalmarsson, H., and Huang, L. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning nonlinear dynamical systems from a single trajectory
Foster, D. J., Rakhlin, A., and Sarkar, T. (2020) · 2004
Cited alongside, same era.
Panel cointegration: asymptotic and finite sample properties of pooled time series tests with an application to the ppp hypothesis
Pedroni, P. (2004) · 2004
Cited alongside, same era.
Least costly identification experiment for control
Bombois, X., Scorletti, G., Gevers, M., Van den Hof, P. M., and Hildebrand, R. (2006) · 2006
Cited alongside, same era.
Dealing with non-stationary environments using context detection
Da Silva, B. C., Basso, E. W., Bazzan, A. L., and Engel, P. M. (2006) · 2006
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R. (2009) · 2009
Cited alongside, same era.
Identification of arx systems with non-stationary inputs—asymptotic analysis with application to adaptive input design
Gerencsér, L., Hjalmarsson, H., and Mårtensson, J. (2009) · 2009
Cited alongside, same era.
System identification of complex and structured systems
Hjalmarsson, H. (2009) · 2009
Cited alongside, same era.
Ouyang, Y., Gagrani, M., and Jain, R. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
Starcraft ii: A new challenge for reinforcement learning
Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A. S., Yeo, M., Makhzani, A., Küttler, H., Agapiou, J., Schrittwieser, J., et al. (2017) · 2017
Later among the works it cites.
Improved regret bounds for thompson sampling in linear quadratic control problems
Abeille, M. and Lazaric, A. (2018) · 2018
Later among the works it cites.
Regret bounds for robust adaptive control of the linear quadratic regulator
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S. (2018) · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S., and Mesbahi, M. (2018) · 2018
Later among the works it cites.
Lecture 20 : Linear dynamics and lqg 3 2 linear system optimal control 2 . 1 linear quadratic regulator ( lqr ) : Discrete-time finite horizon
Jamieson, K., Ademola-Idowu, S. A., and Shi, Y. (2018) · 2018
Later among the works it cites.
Learning-based model predictive control for safe exploration
Koller, T., Berkenkamp, F., Turchetta, M., and Krause, A. (2018) · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Simchowitz, M., Mania, H., Tu, S., Jordan, M. I., and Recht, B. (2018) · 2018
Later among the works it cites.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
Cohen, A., Koren, T., and Mansour, Y. (2019) · 2019
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S. (2019) · 2019
Later among the works it cites.
Certainty equivalence is efficient for linear quadratic control
Mania, H., Tu, S., and Recht, B. (2019) · 2019
Later among the works it cites.
Non-asymptotic identification of lti systems from a single trajectory
Oymak, S. and Ozay, N. (2019) · 2019
Later among the works it cites.
Finite sample system identification: Optimal rates and the role of regularization
Sun, Y., Oymak, S., and Fazel, M. (2020) · 2020
Closest in time.
Online streaming feature selection via multi-conditional independence and mutual information entropy
Wang, H. and You, D. (2020) · 2020
Closest in time.