Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms typically consider discrete-time dynamics, even though the underlying systems are often continuous in time.
An online learning approach to model predictive control
Wagener, N., Cheng, C.-A., Sacks, J., and Boots, B. (2019) · 1902
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Differential dynamic programming
Jacobson, D. H. and Mayne, D. Q. (1970) · 1970
Earlier work this paper cites.
Numerical differentiation and regularization
Cullum, J. (1971) · 1971
Earlier work this paper cites.
Optimal control applied to competing chemotherapeutic cell-kill strategies
Panetta, J. C. and Fister, K. R. (2003) · 1971
Earlier work this paper cites.
A variational method for numerical differentiation
Knowles, I. and Wallace, R. (1995) · 1995
Earlier work this paper cites.
Elements of information theory
Cover, T. M. (1999) · 1999
Earlier work this paper cites.
Model predictive control: past, present and future
Morari, M. and Lee, J. H. (1999) · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K. (2000) · 2000
Earlier work this paper cites.
Ordinary differential equations
Hartman, P. (2002) · 2002
Earlier work this paper cites.
Cranmer, M., Greydanus, S., Hoyer, S., Battaglia, P., Spergel, D., and Ho, S. (2020) · 2003
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
Li, W. and Todorov, E. (2004) · 2004
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Todorov, E. and Li, W. (2005) · 2005
Earlier work this paper cites.
Predicting volatility: getting the most out of return data sampled at different frequencies
Ghysels, E., Santa-Clara, P., and Valkanov, R. (2006) · 2006
Earlier work this paper cites.
Robot modeling and control
Spong, M. W., Hutchinson, S., Vidyasagar, M., et al. (2006) · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Williams, C. K. and Rasmussen, C. E. (2006) · 2006
Earlier work this paper cites.
Heterogeneous multiscale methods: a review
Engquist, B., Li, X., Ren, W., Vanden-Eijnden, E., et al. (2007) · 2007
Earlier work this paper cites.
Optimal outpatient appointment scheduling
Kaandorp, G. C. and Koole, G. (2007) · 2007
Earlier work this paper cites.
Optimal control applied to biological models
Lenhart, S. and Workman, J. T. (2007) · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R. (2008) · 2008
Earlier work this paper cites.
Adaptive optimal control algorithm for continuous-time nonlinear systems based on policy iteration
Vrabie, D. and Lewis, F. L. (2008) · 2008
Earlier work this paper cites.
Differential equations and mathematical biology
Jones, D. S., Plank, M., and Sleeman, B. D. (2009) · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. (2009) · 2009
Earlier work this paper cites.
Online policy iteration based algorithms to solve the continuous-time infinite horizon optimal control problem
Vamvoudakis, K., Vrabie, D., and Lewis, F. (2009) · 2009
Earlier work this paper cites.
Neural network approach to continuous-time direct adaptive optimal control for partially unknown nonlinear systems
Vrabie, D. and Lewis, F. (2009) · 2009
Earlier work this paper cites.
Autonomous flying robots: unmanned aerial vehicles and micro aerial vehicles
Nonami, K., Kendoul, F., Suzuki, S., Wang, W., and Nakazawa, D. (2010) · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Abbasi-Yadkori, Y. and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Numerical differentiation of noisy, nonsmooth data
Chartrand, R. (2011) · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E. (2011) · 2011
Cited alongside, same era.
An introduction to event-triggered and self-triggered control
Heemels, W. P., Johansson, K. H., and Tabuada, P. (2012) · 2012
Cited alongside, same era.
Discovery of complex behaviors through contact-invariant optimization
Mordatch, I., Todorov, E., and Popović, Z. (2012) · 2012
Cited alongside, same era.
Model-based bayesian exploration
Dearden, R., Friedman, N., and Andre, D. (2013) · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B. (2013) · 2013
Cited alongside, same era.
Methods for numerical differentiation of noisy data
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Regularised differentiation of measurement data in systems for monitoring of human movements
Wagner, J., Mazurek, P., Miękina, A., and Morawski, R. Z. (2018) · 2018
Later among the works it cites.
Online learning in kernelized markov decision processes
Chowdhury, S. R. and Gopalan, A. (2019) · 2019
Later among the works it cites.
Hamiltonian neural networks
Greydanus, S., Dzamba, M., and Yosinski, J. (2019) · 2019
Later among the works it cites.
Making deep q-learning methods robust to time discretization
Tallec, C., Blier, L., and Ollivier, Y. (2019) · 2019
Later among the works it cites.
Feedback linearization based on gaussian processes with event-triggered online learning
Umlauft, J. and Hirche, S. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowles, I. and Renka, R. J. (2014) · 2014
Cited alongside, same era.
Linear quadratic tracking control of partially-unknown continuous-time systems using reinforcement learning
Modares, H. and Lewis, F. L. (2014) · 2014
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Osband, I. and Van Roy, B. (2014) · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, D. and Van Roy, B. (2014) · 2014
Cited alongside, same era.
Nonlinear control
Khalil, H. K. (2015) · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Comparison of riemann and lebesgue sampling for first order stochastic systems
Astrom, K. J. and Bernhardsson, B. M. (2002) · 2016
Cited alongside, same era.
Efficient model-based reinforcement learning through optimistic policy search and planning
Curi, S., Berkenkamp, F., and Krause, A. (2020) · 2020
Later among the works it cites.
Model-based reinforcement learning for semi-markov decision processes with neural odes
Du, J., Futoma, J., and Doshi-Velez, F. (2020) · 2020
Later among the works it cites.
Information theoretic regret bounds for online nonlinear control
Kakade, S., Krishnamurthy, A., Lowrey, K., Ohnishi, M., and Sun, W. (2020) · 2020
Later among the works it cites.
New designing approaches for quadcopter using 2d model modelling a cascaded pid controller
Paing, H. S., Schagin, A. V., Win, K. S., and Linn, Y. H. (2020) · 2020
Later among the works it cites.
Naive exploration is optimal for online lqr
Simchowitz, M. and Foster, D. (2020) · 2020
Later among the works it cites.
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., Zariphopoulou, T., and Zhou, X. Y. (2020) · 2020
Later among the works it cites.
Event-triggered and self-triggered control
Heemels, W., Johansson, K. H., and Tabuada, P. (2021) · 2021
Later among the works it cites.
Gaussian process-based real-time learning for safety critical applications
Lederer, A., Conejo, A. J. O., Maier, K. A., Xiao, W., Umlauft, J., and Hirche, S. (2021) · 2021
Later among the works it cites.
Value iteration in continuous actions, states and time
Lutter, M., Mannor, S., Peters, J., Fox, D., and Garg, A. (2021) · 2021
Later among the works it cites.
Convergence and sample complexity of gradient methods for the model-free linear–quadratic regulator problem
Mohammadi, H., Zare, A., Soltanolkotabi, M., and Jovanović, M. R. (2021) · 2021
Later among the works it cites.
Nonlinear sampled-data systems
Nesic, D. and Postoyan, R. (2021) · 2021
Later among the works it cites.
Efficient model-based multi-agent mean-field reinforcement learning
Pasztor, B., Bogunovic, I., and Krause, A. (2021) · 2021
Later among the works it cites.
Distributional gradient matching for learning uncertain neural dynamics models
Treven, L., Wenk, P., Dorfler, F., and Krause, A. (2021) · 2021
Later among the works it cites.
On information gain and regret bounds in gaussian process bandits
Vakili, S., Khezeli, K., and Picheny, V. (2021) · 2021
Later among the works it cites.
Continuous-time model-based reinforcement learning
Yildiz, C., Heinonen, M., and Lähdesmäki, H. (2021) · 2021
Later among the works it cites.
Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon
Basei, M., Guo, X., Hu, A., and Zhang, Y. (2022) · 2022
Later among the works it cites.
A framework and benchmark for deep batch active learning for regression
Holzmüller, D., Zaverkin, V., Kästner, J., and Steinwart, I. (2022) · 2022
Later among the works it cites.
Myriad: a real-world testbed to bridge trajectory optimization and deep learning
Howe, N., Dufort-Labbé, S., Rajkumar, N., and Bacon, P.-L. (2022) · 2022
Later among the works it cites.
Hallucinated adversarial control for conservative offline policy evaluation
Rothfuss, J., Sukhija, B., Birchler, T., Kassraie, P., and Krause, A. (2023) · 2023
Closest in time.
Kinematic bicycle model
Singh, M. and Theers, M. (2021) · 2023
Closest in time.
Model-based causal bayesian optimization
Sussex, S., Makarova, A., and Krause, A. (2023) · 2023
Closest in time.
To sample or not to sample: Self-triggered control for nonlinear systems
Anta, A. and Tabuada, P. (2010) · 2042
Closest in time.