Fetching the paper…
Reading the bibliography…
Learning to control unknown nonlinear dynamical systems is a fundamental problem in reinforcement learning and control theory.
Dual control theory. i
Feldbaum, A. A · 1960
Earlier work this paper cites.
Optimal input signals for parameter estimation in dynamic systems–survey and new results
Mehra, R · 1974
Earlier work this paper cites.
Dynamic system identification: experiment design and data analysis
Goodwin, G. C. and Payne, R. L · 1977
Earlier work this paper cites.
Applied nonlinear control , volume 199
Slotine, J.-J. E., Li, W., et al · 1991
Earlier work this paper cites.
System identification
Ljung, L · 1998
Earlier work this paper cites.
Identification for control: Adaptive input design using convex optimization
Lindqvist, K. and Hjalmarsson, H · 2001
Earlier work this paper cites.
The power of predictions in online control
Yu, C., Shi, G., Chung, S.-J., Yue, Y., and Wierman, A · 2004
Earlier work this paper cites.
Adaptive input design in system identification
Gerencsér, L. and Hjalmarsson, H · 2005
Earlier work this paper cites.
A survey of iterative learning control
Bristow, D. A., Tharayil, M., and Alleyne, A. G · 2006
Earlier work this paper cites.
Adaptive input design for arx systems
Gerencsér, L., Mårtensson, J., and Hjalmarsson, H · 2007
Earlier work this paper cites.
Robust optimal experiment design for system identification
Rojas, C. R., Welsh, J. S., Goodwin, G. C., and Feuer, A · 2007
Earlier work this paper cites.
Uniform approximation of functions with random bases
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Input design for system identification via convex relaxation
Manchester, I. R · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Foundations of modern analysis
Dieudonné, J · 2011
Earlier work this paper cites.
Model learning for robot control: a survey
Nguyen-Tuong, D. and Peters, J · 2011
Earlier work this paper cites.
Application-oriented finite sample experiment design: A semidefinite relaxation approach
Katselis, D., Rojas, C. R., Hjalmarsson, H., and Bengtsson, M · 2012
Earlier work this paper cites.
Adaptive control
Åström, K. J. and Wittenmark, B · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Osband, I. and Van Roy, B · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Information theoretic mpc for model-based reinforcement learning
Williams, G., Wagener, N., Goldfain, B., Drews, P., Rehg, J. M., Boots, B., and Theodorou, E. A · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Stochastic model predictive control with active uncertainty learning: A survey on dual control
Mesbah, A · 2018
Cited alongside, same era.
Information theoretic regret bounds for online nonlinear control
Kakade, S., Krishnamurthy, A., Lowrey, K., Ohnishi, M., and Sun, W · 2020
Later among the works it cites.
Chance-constrained trajectory optimization for safe exploration and learning of nonlinear systems
Nakka, Y. K., Liu, A., Shi, G., Anandkumar, A., Yue, Y., and Chung, S.-J · 2020
Later among the works it cites.
Naive exploration is optimal for online lqr
Simchowitz, M. and Foster, D · 2020
Later among the works it cites.
Improper learning for non-stochastic control
Simchowitz, M., Singh, K., and Hazan, E · 2020
Later among the works it cites.
Active learning for identification of linear dynamical systems
Wagenmaker, A. and Jamieson, K · 2020
Later among the works it cites.
Regret bounds for adaptive nonlinear control
Boffi, N. M., Tu, S., and Slotine, J.-J. E · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simchowitz, M., Mania, H., Tu, S., Jordan, M. I., and Recht, B · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
Cohen, A., Koren, T., and Mansour, Y · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Cited alongside, same era.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Cited alongside, same era.
Certainty equivalence is efficient for linear quadratic control
Mania, H., Tu, S., and Recht, B · 2019
Cited alongside, same era.
Stochastic gradient descent learns state equations with nonlinear activations
Oymak, S · 2019
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
Du, S., Kakade, S., Lee, J., Lovett, S., Mahajan, G., Sun, W., and Wang, R · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Foster, D. J., Kakade, S. M., Qian, J., and Rakhlin, A · 2021
Later among the works it cites.
Adaptive-control-oriented meta-learning for nonlinear systems
Richards, S. M., Azizan, N., Slotine, J.-J., and Pavone, M · 2021
Later among the works it cites.
Pc-mlp: Model-based reinforcement learning with policy cover guided exploration
Song, Y. and Sun, W · 2021
Later among the works it cites.
Task-optimal exploration in linear dynamical systems
Wagenmaker, A. J., Simchowitz, M., and Jamieson, K · 2021
Later among the works it cites.
Reward is enough for convex mdps
Zahavy, T., O’Donoghue, B., Desjardins, G., and Singh, S · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D., Gu, Q., and Szepesvari, C · 2021
Later among the works it cites.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Brunke, L., Greeff, M., Hall, A. W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A. P · 2022
Later among the works it cites.
Active learning for nonlinear system identification with guarantees
Mania, H., Jordan, M. I., and Recht, B · 2022
Later among the works it cites.
Neural-fly enables rapid learning for agile flight in strong winds
O’Connell, M., Shi, G., Shi, X., Azizzadenesheli, K., Anandkumar, A., Yue, Y., and Chung, S.-J · 2022
Later among the works it cites.
Non-asymptotic and accurate learning of nonlinear dynamical systems
Sattar, Y. and Oymak, S · 2022
Later among the works it cites.
Instance-dependent near-optimal policy identification in linear mdps via online experiment design
Wagenmaker, A. and Jamieson, K · 2022
Later among the works it cites.
Leveraging offline data in online reinforcement learning
Wagenmaker, A. and Pacchiano, A · 2022
Later among the works it cites.