Fetching the paper…
Reading the bibliography…
The deployment of Reinforcement Learning (RL) in real-world applications is constrained by its failure to satisfy safety criteria.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 1912
Earlier work this paper cites.
An augmented lagrangian treatment of contact problems involving friction
J Ci Simo and TA1143885 Laursen · 1992
Earlier work this paper cites.
Constrained Markov decision processes: stochastic modeling
Eitan Altman · 1999
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Model predictive controllers
Eduardo F Camacho, Carlos Bordons, Eduardo F Camacho, and Carlos Bordons · 2007
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
A bi-symmetric log transformation for wide-range data
J Beau W Webber · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Survey of model-based reinforcement learning: Applications on robotics
Athanasios S Polydoros and Lazaros Nalpantidis · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Learning-based model predictive control for safe exploration
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause · 2018
Earlier work this paper cites.
Constrained cross-entropy method for safe reinforcement learning
Min Wen and Ufuk Topcu · 2018
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Cited alongside, same era.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Cited alongside, same era.
Constrained model-based reinforcement learning with robust cross-entropy method
Zuxin Liu, Hongyi Zhou, Baiming Chen, Sicheng Zhong, Martial Hebert, and Ding Zhao · 2020
Model-based safe deep reinforcement learning via a constrained proximal policy optimization algorithm
Ashish K Jayant and Shalabh Bhatnagar · 2022
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun · 2022
Later among the works it cites.
Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou · 2022
Later among the works it cites.
Learning off-policy with online planning
Harshit Sikchi, Wenxuan Zhou, and David Held · 2022
Later among the works it cites.
Constrained update projection approach to safe policy optimization
Long Yang, Jiaming Ji, Juntao Dai, Linrui Zhang, Binbin Zhou, Pengfei Li, Yaodong Yang, and Gang Pan · 2022
Later among the works it cites.
Augmented proximal policy optimization for safe reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Cited alongside, same era.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Cited alongside, same era.
Safe reinforcement learning by imagining the near future
Garrett Thomas, Yuping Luo, and Tengyu Ma · 2021
Cited alongside, same era.
A predictive safety filter for learning-based control of constrained nonlinear dynamical systems, 2021
Kim P. Wabersich and Melanie N. Zeilinger · 2021
Cited alongside, same era.
Constrained policy optimization via bayesian world models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause · 2022
Cited alongside, same era.
Juntao Dai, Jiaming Ji, Long Yang, Qian Zheng, and Gang Pan · 2023
Closest in time.
Dense reinforcement learning for safety validation of autonomous vehicles
Shuo Feng, Haowei Sun, Xintao Yan, Haojie Zhu, Zhengxia Zou, Shengyin Shen, and Henry X Liu · 2023
Closest in time.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.
Autocost: Evolving intrinsic cost for zero-violation reinforcement learning
Tairan He, Weiye Zhao, and Changliu Liu · 2023
Closest in time.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, Aske Plaat, Catholijn M Jonker, et al · 2023
Closest in time.
Rosarl: Reward-only safe reinforcement learning
Geraud Nangue Tasse, Tamlin Love, Mark Nemecek, Steven James, and Benjamin Rosman · 2023
Closest in time.
Safe trajectory sampling in model-based reinforcement learning
Sicelukwanda Zwane, Denis Hadjivelichkov, Yicheng Luo, Yasemin Bekiroglu, Dimitrios Kanoulas, and Marc Peter Deisenroth · 2023
Closest in time.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Voce: Variational optimization with conservative estimation for offline safe reinforcement learning
Jiayi Guan, Guang Chen, Jiaming Ji, Long Yang, Zhijun Li, et al · 2024
Closest in time.
Fedfed: Feature distillation against data heterogeneity in federated learning
Zhiqin Yang, Yonggang Zhang, Yu Zheng, Xinmei Tian, Hao Peng, Tongliang Liu, and Bo Han · 2024
Closest in time.