Fetching the paper…
Reading the bibliography…
Improving sample-efficiency and safety are crucial challenges when deploying reinforcement learning in high-stakes real world applications.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy P. Lillicrap, Jimmy Ba, and Mohammad Norouzi · 1912
Earlier work this paper cites.
Constrained Markov Decision Processes
E. Altman · 1999
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Robust Control of Markov Decision Processes with Uncertain Transition Matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2006
Earlier work this paper cites.
The cross-entropy method for continuous multi-extremal optimization
Dirk P. Kroese, Sergey Porotsky, and Reuven Y. Rubinstein · 2006
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
A stochastic approximation method
Herbert E. Robbins · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Peter Deisenroth and Carl Edward Rasmussen · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Scaling up robust mdps by reinforcement learning, 2013
Aviv Tamar, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Auto-encoding variational bayes, 2014
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier García, Fern, and o Fernández · 2015
Cited alongside, same era.
Concrete problems in ai safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus), 2016
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Cited alongside, same era.
Constrained policy optimization, 2017
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees, 2017
Soft actor-critic algorithms and applications, 2019
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2019
Later among the works it cites.
Averaging weights leads to wider optima and better generalization, 2019
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization, 2019
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
A simple baseline for bayesian uncertainty in deep learning, 2019
Wesley Maddox, Timur Garipov, Pavel Izmailov, Dmitry Vetrov, and Andrew Gordon Wilson · 2019
Later among the works it cites.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Felix Berkenkamp, Matteo Turchetta, Angela P. Schoellig, and Andreas Krause · 2017
Cited alongside, same era.
Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning, 2017
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft · 2017
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning, 2017
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles, 2017
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces, 2018
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
Chen Tessler, Yonathan Efroni, and Shie Mannor · 2019
Later among the works it cites.
Exploration-exploitation in constrained mdps, 2020
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Later among the works it cites.
Monte carlo gradient estimation in machine learning, 2020
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods, 2020
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Conservative safety critics for exploration, 2021
Homanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine, Florian Shkurti, and Animesh Garg · 2021
Later among the works it cites.
Combining pessimism with optimism for robust and efficient model-based deep reinforcement learning, 2021
Sebastian Curi, Ilija Bogunovic, and Andreas Krause · 2021
Later among the works it cites.
Augmented lagrangian method for instantaneously constrained reinforcement learning problems
Jingqi Li, David Fridovich-Keil, Somayeh Sojoudi, and Claire J. Tomlin · 2021
Later among the works it cites.
Constrained model-based reinforcement learning with robust cross-entropy method, 2021
Zuxin Liu, Hongyi Zhou, Baiming Chen, Sicheng Zhong, Martial Hebert, and Ding Zhao · 2021
Later among the works it cites.
Safe reinforcement learning via curriculum induction, 2021
Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, and Alekh Agarwal · 2021
Later among the works it cites.
A predictive safety filter for learning-based control of constrained nonlinear dynamical systems, 2021
Kim P. Wabersich and Melanie N. Zeilinger · 2021
Later among the works it cites.
Safe continuous control with constrained model-based policy optimization, 2021
Moritz A. Zanger, Karam Daaboul, and J. Marius Zöllner · 2021
Later among the works it cites.