Fetching the paper…
Reading the bibliography…
An important facet of reinforcement learning (RL) has to do with how the agent goes about exploring the environment.
Risk-sensitive Markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
TD algorithm for the variance of return and mean-variance reinforcement learning
Makoto Sato, Hajime Kimura, and Shibenobu Kobayashi · 2001
Earlier work this paper cites.
Stability theory of dynamical systems
Nam Parshad Bhatia and Giorgio P Szegö · 2002
Earlier work this paper cites.
Sparse on-line Gaussian processes
Lehel Csató and Manfred Opper · 2002
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2004
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2005
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Peter Geibel and Fritz Wysotzki · 2005
Earlier work this paper cites.
Risk-directed Exploration in Reinforcement Learning
Edith LM Law, Melanie Coggan, Doina Precup, and Bohdana Ratitch · 2005
Earlier work this paper cites.
A hybrid reinforcement learning approach to autonomic resource allocation
Gerald Tesauro, Nicholas K Jong, Rajarshi Das, and Mohamed N Bennani · 2006
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Bayesian approach to global optimization: theory and applications , volume 37
Jonas Mockus · 2012
Earlier work this paper cites.
Safe exploration in Markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Cited alongside, same era.
Policy Gradients with Variance Related Risk Criteria
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Help an agent out: Student/teacher learning in sequential decision tasks
Lisa Torrey and Matthew E Taylor · 2012
Cited alongside, same era.
Smart exploration in reinforcement learning using absolute temporal difference errors
Clement Gehring and Doina Precup · 2013
Cited alongside, same era.
Off-policy reinforcement learning with Gaussian processes
Girish Chowdhary, Miao Liu, Robert Grande, Thomas Walsh, Jonathan How, and Lawrence Carin · 2014
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Later among the works it cites.
On kernelized multi-armed bandits
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Later among the works it cites.
Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Safe exploration techniques for reinforcement learning–an overview
Martin Pecka and Tomas Svoboda · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Safe learning of regions of attraction for uncertain, nonlinear systems with gaussian processes
Felix Berkenkamp, Riccardo Moriconi, Angela P Schoellig, and Andreas Krause · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Večerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Later among the works it cites.
Safe reinforcement learning via shielding
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu · 2018
Later among the works it cites.
A Lyapunov-based Approach to Safe Reinforcement Learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.