Fetching the paper…
Reading the bibliography…
Reinforcement learning is a promising approach to synthesizing policies for challenging robotics tasks.
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Neuronlike adaptive elements that can solve difficult learning control problems,” IEEE transactions on systems, man, and cybernetics , no. 5, pp. 834–846, 1983
1983
Earlier work this paper cites.
P. A. Parrilo, “Structured semidefinite programs and semialgebraic geometry methods in robustness and optimization,” Ph.D. dissertation, California Institute of Technology, 2000
2000
Earlier work this paper cites.
T. J. Perkins and A. G. Barto, “Lyapunov design for safe reinforcement learning,” JMLR , 2002
2002
Earlier work this paper cites.
R. Tedrake, “Lqr-trees: Feedback motion planning on sparse randomized trees,” in RSS , 2009
2009
Earlier work this paper cites.
J. H. Gillula and C. J. Tomlin, “Guaranteed safe online learning via reachability: tracking a ground target using a quadrotor,” in ICRA , 2012
2012
Earlier work this paper cites.
T. M. Moldovan and P. Abbeel, “Safe exploration in markov decision processes,” in ICML , 2012
2012
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning , 2013, pp. 1–9
2013
Earlier work this paper cites.
L. Asselborn, D. Gross, and O. Stursberg, “Control of uncertain nonlinear systems using ellipsoidal reachability calculus,” IFAC Proceedings Volumes , vol. 46, no. 23, pp. 50–55, 2013
2013
Earlier work this paper cites.
A. K. Akametalu, S. Kaynama, J. F. Fisac, M. N. Zeilinger, J. H. Gillula, and C. J. Tomlin, “Reachability-based safe learning with gaussian processes,” in CDC . Citeseer, 2014, pp. 1424–1431
2014
Earlier work this paper cites.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” JMLR , 2015
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
M. Turchetta, F. Berkenkamp, and A. Krause, “Safe exploration in finite markov decision processes with gaussian processes,” in NIPS , 2016
2016
Cited alongside, same era.
Y. Wu, R. Shariff, T. Lattimore, and C. Szepesvári, “Conservative bandits,” in ICML , 2016
2016
Cited alongside, same era.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in ICML , 2017
2017
Cited alongside, same era.
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause, “Safe model-based reinforcement learning with stability guarantees,” in Advances in Neural Information Processing Systems , 2017, pp. 908–918
2017
M. Wen and U. Topcu, “Constrained cross-entropy method for safe reinforcement learning,” in Advances in Neural Information Processing Systems , 2018, pp. 7450–7460
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Koller, F. Berkenkamp, M. Turchetta, and A. Krause, “Learning-based model predictive control for safe exploration,” in 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 2018, pp. 6059–6066
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2018
Cited alongside, same era.
O. Bastani, P. Yewen, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” in Advances in neural information processing systems , 2018
2018
Cited alongside, same era.
Y. Chow, O. Nachum, and E. Duenez-Guzman, “A lyapunov-based approach to safe reinforcement learning,” in NeurIPS , 2018
2018
Cited alongside, same era.
M. Alshiekh, R. Bloem, R. Ehlers, B. Konighofer, S. Niekum, and U. Topcu, “Safe reinforcement learning via shielding,” in AAAI , 2018
2018
Cited alongside, same era.
K. P. Wabersich and M. N. Zeilinger, “Linear model predictive safety certification for learning-based control,” in 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 2018, pp. 7130–7135
2018
Cited alongside, same era.
——, Underactuated Robotics: Algorithms for Walking, Running, Swimming, Flying, and Manipulation , 2018. [Online]. Available: http://underactuated.mit.edu/
2018
Later among the works it cites.
A. Khan, E. Tolstaya, A. Ribeiro, and V. Kumar, “Graph policy gradients for large scale robot control,” in CORL , 2019
2019
Closest in time.
R. Ivanov, J. Weimer, R. Alur, G. J. Pappas, and I. Lee, “Verisig: verifying safety properties of hybrid systems with neural network controllers,” in HSCC , 2019
2019
Closest in time.
H. Zhu, Z. Xiong, S. Magill, and S. Jagannathan, “An inductive synthesis framework for verifiable reinforcement learning,” in PLDI , 2019
2019
Closest in time.
2019
Closest in time.
O. Bastani, “Sample complexity of estimating the policy gradient for nearly deterministic dynamical systems,” in AISTATS , 2020
2020
Closest in time.