Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is promising for complicated stochastic nonlinear control problems.
Krishna Prakashan Media, 1968
H. L. Royden, Real analysis · 1968
Earlier work this paper cites.
M. Corless and G. Leitmann, “Continuous state feedback guaranteeing uniform ultimate boundedness for uncertain dynamic systems,” IEEE Transactions on Automatic Control
1981
Earlier work this paper cites.
A. Thowsen, “Uniform ultimate boundedness of the solutions of uncertain dynamic delay systems with state-dependent and memoryless feedback control,” International Journal of control
1983
Earlier work this paper cites.
C. E. Garcia, D. M. Prett, and M. Morari, “Model predictive control: theory and practice—a survey,” Automatica
1989
Earlier work this paper cites.
Prentice hall Englewood Cliffs, NJ, 1991
J.-J. E. Slotine, W. Li, et al · 1991
Earlier work this paper cites.
R. S. Sutton, A. G. Barto, and R. J. Williams, “Reinforcement learning is direct adaptive optimal control,” IEEE Control Systems Magazine
1992
Earlier work this paper cites.
D. P. Bertsekas, “Nonlinear programming,” Journal of the Operational Research Society
1997
Earlier work this paper cites.
CRC Press, 1999
E. Altman, Constrained Markov decision processes · 1999
Earlier work this paper cites.
D. Q. Mayne, J. B. Rawlings, C. V. Rao, and P. O. Scokaert, “Constrained model predictive control: Stability and optimality,” Automatica
2000
Earlier work this paper cites.
E. Boukas and Z. Liu, “Robust stability and h/sub/spl infin//control of discrete-time jump linear systems with time-delay: an lmi approach,” in Decision and Control, 2000. Proceedings of the 39th IEEE Conference on
2000
Earlier work this paper cites.
D. Q. Mayne, “Control of constrained dynamic systems,” European Journal of Control
2001
Earlier work this paper cites.
Siam, 2002
M. Vidyasagar, Nonlinear systems analysis · 2002
Earlier work this paper cites.
P. Shih, B. Kaul, S. Jagannathan, and J. Drallmeier, “Near optimal output-feedback control of nonlinear discrete-time systems in nonstrict feedback form with application to engines,” in 2007 International Joint Conference on Neural Networks
2007
Earlier work this paper cites.
R. Scattolini, “Architectures for distributed and hierarchical model predictive control–a review,” Journal of process control
2009
Earlier work this paper cites.
R. S. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in neural information processing systems
2009
Earlier work this paper cites.
J. Huang, Z. Han, X. Cai, and L. Liu, “Uniformly ultimately bounded tracking control of linear differential inclusions with stochastic disturbance,” Mathematics and Computers in Simulation
2011
Earlier work this paper cites.
T. M. Moldovan and P. Abbeel, “Safe exploration in markov decision processes,” ICML
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
2012
Earlier work this paper cites.
Springer Science & Business Media, 2013
S. Sastry, Nonlinear systems: analysis, stability, and control · 2013
Cited alongside, same era.
N. Korda and P. La, “On td (0) with function approximation: Concentration bounds and a centered variant with exponential convergence,” in International Conference on Machine Learning
2015
Cited alongside, same era.
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning
2015
Cited alongside, same era.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research
2015
W. Saunders, G. Sastry, A. Stuhlmueller, and O. Evans, “Trial without error: Towards safe reinforcement learning via human intervention,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
X. Yang, D. Liu, Q. Wei, and D. Wang, “Guaranteed cost neural tracking control for a class of uncertain nonlinear systems using adaptive dynamic programming,” Neurocomputing
2016
Cited alongside, same era.
C. Mu, Z. Ni, C. Sun, and H. He, “Data-driven tracking control with adaptive dynamic programming for a class of continuous-time nonlinear systems,” IEEE transactions on cybernetics
2016
Cited alongside, same era.
C. J. Ostafew, A. P. Schoellig, and T. D. Barfoot, “Robust constrained learning-based nmpc enabling reliable mobile robot path tracking,” The International Journal of Robotics Research
2016
Cited alongside, same era.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Thirtieth AAAI conference on artificial intelligence
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. K. Jain and S. Bhasin, “Uniformly ultimately bounded tracking for uncertain euler-lagrange systems with unknown time-varying input delay,” IFAC-PapersOnLine
2017
Cited alongside, same era.
S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International Conference on Machine Learning
2018
Later among the works it cites.
2018
Later among the works it cites.
B. Luo, Y. Yang, D. Liu, and H.-N. Wu, “Event-triggered optimal control with performance guarantees using adaptive dynamic programming,” IEEE transactions on neural networks and learning systems
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Zou, T. Xu, and Y. Liang, “Finite-sample analysis for sarsa with linear function approximation,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
W. Yu, V. C. Kumar, G. Turk, and C. K. Liu, “Sim-to-real transfer for biped locomotion,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
B. Thananjeyan, A. Balakrishna, U. Rosolia, F. Li, R. McAllister, J. E. Gonzalez, S. Levine, F. Borrelli, and K. Goldberg, “Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks,” IEEE Robotics and Automation Letters
2020
Closest in time.
M. Zanon and S. Gros, “Safe reinforcement learning using robust mpc,” IEEE Transactions on Automatic Control
2020
Closest in time.
2020
Closest in time.
J. Harrison, A. Garg, B. Ivanovic, Y. Zhu, S. Savarese, L. Fei-Fei, and M. Pavone, “Adapt: zero-shot adaptive policy transfer for stochastic dynamical systems,” in Robotics Research
2020
Closest in time.