Fetching the paper…
Reading the bibliography…
We address the problem of safe reinforcement learning from pixel observations.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes: Stochastic Modeling
Eitan Altman · 1999
Earlier work this paper cites.
Actor-Critic Algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
Exploration and Apprenticeship Learning in Reinforcement Learning
Pieter Abbeel and Andrew Y. Ng · 2005
Earlier work this paper cites.
Piecewise Linear Dynamic Programming for Constrained POMDPs
Joshua D. Isom, Sean P. Meyn, and Richard D. Braatz · 2008
Earlier work this paper cites.
Point-Based Value Iteration for Constrained POMDPs
Dongho Kim, Jaesong Lee, Kee-Eung Kim, and Pascal Poupart · 2011
Earlier work this paper cites.
Safe Exploration Techniques for Reinforcement Learning - An Overview
Martin Pecka and Tomás Svoboda · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust Region Policy Optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
A Comprehensive Survey on Safe Reinforcement Learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Constrained Policy Optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Safe Reinforcement Learning via Shielding
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Monte-Carlo Tree Search for Constrained POMDPs
Jongmin Lee, Geon-hyeong Kim, Pascal Poupart, and Kee-Eung Kim · 2018
Cited alongside, same era.
Column Generation Algorithms for Constrained POMDPs
Erwin Walraven and Matthijs T. J. Spaan · 2018
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
Evaluating strategic structures in multi-agent inverse reinforcement learning
Justin Fu, Andrea Tacchetti, Julien Perolat, and Yoram Bachrach · 2021
Later among the works it cites.
Policy Learning with Constraints in Model-free Reinforcement Learning: A Survey
Yongshuai Liu, Avishai Halev, and Xin Liu · 2021
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Gabriel Dulac-Arnold, Nir Levine, Daniel J. Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2021
Later among the works it cites.
Safe Continuous Control with Constrained Model-Based Policy Optimization
Moritz A. Zanger, Karam Daaboul, and J. Marius Zöllner · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Benchmarking Safe Exploration in Deep Reinforcement Learning, 2019
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Learning Latent Dynamics for Planning from Pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Cited alongside, same era.
Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model
Alex X. Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Cited alongside, same era.
Safe Reinforcement Learning Using Probabilistic Shields (Invited Paper)
Nils Jansen, Bettina Könighofer, Sebastian Junges, Alex Serban, and Roderick Bloem · 2020
Cited alongside, same era.
Learning to Walk in the Real World with Minimal Human Effort
Sehoon Ha, Peng Xu, Zhenyu Tan, Sergey Levine, and Jie Tan · 2020
Cited alongside, same era.
WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement Learning
Qisong Yang, Thiago D Simão, Simon H Tindemans, and Matthijs TJ Spaan · 2021
Later among the works it cites.
Constrained Policy Optimization via Bayesian World Models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause · 2022
Closest in time.
Safe Reinforcement Learning via Shielding under Partial Observability
Steven Carr, Nils Jansen, Sebastian Junges, and Ufuk Topcu · 2022
Closest in time.
Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)
Peter Vamplew, Benjamin J. Smith, Johan Källström, Gabriel Ramos, Roxana Rădulescu, Diederik M. Roijers, Conor F. Hayes, Fredrik Heintz, Patrick Mannion, Pieter J. K. Libin, Richard Dazeley, and Cameron Foale · 2022
Closest in time.
Direct Behavior Specification via Constrained Reinforcement Learning
Julien Roy, Roger Girgis, Joshua Romoff, Pierre-Luc Bacon, and Chris J Pal · 2022
Closest in time.
Maximum Entropy RL (Provably) Solves Some Robust RL Problems
Benjamin Eysenbach and Sergey Levine · 2022
Closest in time.
Safety-constrained reinforcement learning with a distributional safety critic
Qisong Yang, Thiago D Simão, Simon H Tindemans, and Matthijs TJ Spaan · 2022
Closest in time.