Fetching the paper…
Reading the bibliography…
Exploration in unknown environments is a fundamental problem in reinforcement learning and control.
Dynamic programming under uncertainty with a quadratic criterion function
Simon, H. A · 1956
Earlier work this paper cites.
A note on certainty equivalence in dynamic planning
Theil, H · 1957
Earlier work this paper cites.
Synthesis of optimal inputs for multiinput-multioutput (mimo) systems with process noise part i: Frequenc y-domain synthesis part ii: Time-domain synthesis
Mehra, R. K · 1976
Earlier work this paper cites.
Dynamic system identification: experiment design and data analysis
Goodwin, G. C. and Payne, R. L · 1977
Earlier work this paper cites.
Nonlinear experiments: Optimal design and inference based on likelihood
Chaudhuri, P. and Mykland, P. A · 1993
Earlier work this paper cites.
Applications of the van trees inequality: a bayesian cramér-rao bound
Gill, R. D., Levit, B. Y., et al · 1995
Earlier work this paper cites.
For model-based control design, closed-loop identification gives better performance
Hjalmarsson, H., Gevers, M., and De Bruyne, F · 1996
Earlier work this paper cites.
Identification for control: Adaptive input design using convex optimization
Lindqvist, K. and Hjalmarsson, H · 2001
Earlier work this paper cites.
Identification for control: optimal input design with respect to a worst-case ν \nu -gap cost function
Hildebrand, R. and Gevers, M · 2002
Earlier work this paper cites.
Applications of mixed h2 and hinfin; input design in identification
Barenthin, M., Jansson, H., and Hjalmarsson, H · 2005
Earlier work this paper cites.
Adaptive input design in system identification
Gerencsér, L. and Hjalmarsson, H · 2005
Earlier work this paper cites.
Input design via lmis admitting frequency-wise model specifications in confidence regions
Jansson, H. and Hjalmarsson, H · 2005
Earlier work this paper cites.
Optimal design of experiments
Pukelsheim, F · 2006
Earlier work this paper cites.
Adaptive input design for arx systems
Gerencsér, L., Mårtensson, J., and Hjalmarsson, H · 2007
Earlier work this paper cites.
Robust optimal experiment design for system identification
Rojas, C. R., Welsh, J. S., Goodwin, G. C., and Feuer, A · 2007
Earlier work this paper cites.
Identification of arx systems with non-stationary inputs—asymptotic analysis with application to adaptive input design
Gerencsér, L., Hjalmarsson, H., and Mårtensson, J · 2009
Earlier work this paper cites.
Identification and the information matrix: how to get just sufficiently rich?
Gevers, M., Bazanella, A. S., Bombois, X., and Miskovic, L · 2009
Earlier work this paper cites.
Input design for system identification via convex relaxation
Manchester, I. R · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Abbasi-Yadkori, Y. and Szepesvári, C · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Robustness in experiment design
Rojas, C. R., Aguero, J.-C., Welsh, J. S., Goodwin, G. C., and Feuer, A · 2011
Earlier work this paper cites.
On the fundamental limits of adaptive sensing
Arias-Castro, E., Candes, E. J., and Davenport, M. A · 2012
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Hsu, D., Kakade, S., Zhang, T., et al · 2012
Cited alongside, same era.
Application-oriented finite sample experiment design: A semidefinite relaxation approach
Katselis, D., Rojas, C. R., Hjalmarsson, H., and Bengtsson, M · 2012
Cited alongside, same era.
Robust input design for resonant systems under limited a priori information
Larsson, C., Geerardyn, E., and Schoukens, J · 2012
Cited alongside, same era.
Adaptive control
Åström, K. J. and Wittenmark, B · 2013
Cited alongside, same era.
Robust and adaptive excitation signal generation for input and output constrained systems
Hägg, P., Larsson, C. A., and Hjalmarsson, H · 2013
Cited alongside, same era.
Design of experiments in nonlinear models
Pronzato, L. and Pázman, A · 2013
Cited alongside, same era.
System level synthesis
Anderson, J., Doyle, J. C., Low, S. H., and Matni, N · 2019
Later among the works it cites.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
Cohen, A., Koren, T., and Mansour, Y · 2019
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E · 2019
Later among the works it cites.
Safely learning to control the constrained linear quadratic regulator
Dean, S., Tu, S., Matni, N., and Recht, B · 2019
Later among the works it cites.
Sample complexity lower bounds for linear system identification
Jedra, Y. and Proutiere, A · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence rates of active learning for maximum likelihood estimation
Chaudhuri, K., Kakade, S., Netrapalli, P., and Sanghavi, S · 2015
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Kaufmann, E., Cappé, O., and Garivier, A · 2016
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S · 2017
Cited alongside, same era.
The simulator: Understanding adaptive sampling in the moderate-confidence regime
Simchowitz, M., Jamieson, K., and Recht, B · 2017
Cited alongside, same era.
Mania, H., Tu, S., and Recht, B · 2019
Later among the works it cites.
Non-asymptotic identification of lti systems from a single trajectory
Oymak, S. and Ozay, N · 2019
Later among the works it cites.
Near optimal finite time identification of arbitrary linear dynamical systems
Sarkar, T. and Rakhlin, A · 2019
Later among the works it cites.
Finite-time system identification for partially observed lti systems of unknown order
Sarkar, T., Rakhlin, A., and Dahleh, M. A · 2019
Later among the works it cites.
Learning linear dynamical systems with semi-parametric least squares
Simchowitz, M., Boczar, R., and Recht, B · 2019
Later among the works it cites.
Finite sample analysis of stochastic system identification
Tsiamis, A. and Pappas, G. J · 2019
Later among the works it cites.
Almost horizon-free structure-aware best policy identification with a generative model
Zanette, A., Kochenderfer, M., and Brunskill, E · 2019
Later among the works it cites.
Efficient optimistic exploration in linear-quadratic regulators via lagrangian relaxation
Abeille, M. and Lazaric, A · 2020
Later among the works it cites.
Information theoretic regret bounds for online nonlinear control
Kakade, S., Krishnamurthy, A., Lowrey, K., Ohnishi, M., and Sun, W · 2020
Later among the works it cites.
Active learning for nonlinear system identification with guarantees
Mania, H., Jordan, M. I., and Recht, B · 2020
Later among the works it cites.
Best policy identification in discounted mdps: Problem-specific sample complexity
Marjani, A. A. and Proutiere, A · 2020
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Ménard, P., Domingues, O. D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M · 2020
Later among the works it cites.
Naive exploration is optimal for online lqr
Simchowitz, M. and Foster, D. J · 2020
Later among the works it cites.
Improper learning for non-stochastic control
Simchowitz, M., Singh, K., and Hazan, E · 2020
Later among the works it cites.
Active learning for identification of linear dynamical systems
Wagenmaker, A. and Jamieson, K · 2020
Later among the works it cites.
On uninformative optimal policies in adaptive lqr with unknown b-matrix
Ziemann, I. and Sandberg, H · 2020
Later among the works it cites.
Navigating to the best policy in markov decision processes
Marjani, A. A., Garivier, A., and Proutiere, A · 2021
Closest in time.