Fetching the paper…
Reading the bibliography…
Reinforcement learning is a powerful paradigm for learning optimal policies from experimental data.
Bayesian approach to global optimization
Jonas Mockus · 1989
Earlier work this paper cites.
Nonlinear systems
Hassan K. Khalil and J. W. Grizzle · 1996
Earlier work this paper cites.
Multidimensional triangulation and interpolation for reinforcement learning
Scott Davies · 1996
Earlier work this paper cites.
Reinforcement learning: an introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Stefano P. Coraluppi and Steven I. Marcus · 1999
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf · 2002
Earlier work this paper cites.
Lyapunov design for safe reinforcement learning
Theodore J. Perkins and Andrew G. Barto · 2003
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Peter Geibel and Fritz Wysotzki · 2005
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Carl Edward Rasmussen and Christopher K.I Williams · 2006
Earlier work this paper cites.
Approximate dynamic programming: solving the curses of dimensionality
Warren B. Powell · 2007
Earlier work this paper cites.
Safe exploration for reinforcement learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, and Steffen Udluft · 2008
Earlier work this paper cites.
Support Vector Machines
Andreas Christmann and Ingo Steinwart · 2008
Earlier work this paper cites.
Robust Markov Decision Processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2012
Earlier work this paper cites.
Safe exploration in Markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Cited alongside, same era.
Safe exploration of state and action spaces in reinforcement learning
J. Garcia and F. Fernandez · 2012
Cited alongside, same era.
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Niranjan Srinivas, Andreas Krause, Sham M. Kakade, and Matthias Seeger · 2012
Cited alongside, same era.
Provably safe and robust learning-based model predictive control
Anil Aswani, Humberto Gonzalez, S. Shankar Sastry, and Claire Tomlin · 2013
Cited alongside, same era.
Safe exploration techniques for reinforcement learning – an overview
Martin Pecka and Tomas Svoboda · 2014
Cited alongside, same era.
Scaling Up Robust MDPs by Reinforcement Learning
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Later among the works it cites.
Safe exploration in finite markov decision processes with gaussian processes
Matteo Turchetta, Felix Berkenkamp, and Andreas Krause · 2016
Later among the works it cites.
Safe controller optimization for quadrotors with Gaussian processes
Felix Berkenkamp, Angela P. Schoellig, and Andreas Krause · 2016
Later among the works it cites.
Safe control under uncertainty with Probabilistic Signal Temporal Logic
Dorsa Sadigh and Ashish Kapoor · 2016
Later among the works it cites.
Robust constrained learning-based NMPC enabling reliable mobile robot path tracking
Chris J. Ostafew, Angela P. Schoellig, and Timothy D. Barfoot · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reachability-based safe learning with Gaussian processes
Anayo K. Akametalu, Shahab Kaynama, Jaime F. Fisac, Melanie N. Zeilinger, Jeremy H. Gillula, and Claire J. Tomlin · 2014
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Cited alongside, same era.
Safe exploration for active learning with Gaussian processes
Jens Schreiter, Duy Nguyen-Tuong, Mona Eberts, Bastian Bischoff, Heiner Markert, and Marc Toussaint · 2015
Cited alongside, same era.
Safe exploration for optimization with Gaussian processes
Yanan Sui, Alkis Gotovos, Joel W. Burdick, and Andreas Krause · 2015
Cited alongside, same era.
A sampling approach to finding Lyapunov functions for nonlinear discrete-time systems
Ruxandra Bobiti and Mircea Lazar · 2016
Later among the works it cites.
Safe learning of regions of attraction in nonlinear systems with Gaussian processes
Felix Berkenkamp, Riccardo Moriconi, Angela P. Schoellig, and Andreas Krause · 2016
Later among the works it cites.
Stability of controllers for Gaussian process forward models
Julia Vinogradska, Bastian Bischoff, Duy Nguyen-Tuong, Henner Schmidt, Anne Romer, and Jan Peters · 2016
Later among the works it cites.
Computation of local ISS Lyapunov functions for discrete-time systems via linear programming
Huijuan Li and Lars Grüne · 2016
Later among the works it cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Closest in time.
On kernelized multi-armed bandits
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Closest in time.
GPflow: a Gaussian process library using TensorFlow
Alexander G. de G. Matthews, Mark van der Wilk, Tom Nickson, Keisuke Fujii, Alexis Boukouvalas, Pablo León-Villagrá, Zoubin Ghahramani, and James Hensman · 2017
Closest in time.