Fetching the paper…
Reading the bibliography…
We address the issue of safety in reinforcement learning.
Adaptive control of Markov chains, I: Finite parameter set
V. Borkar and P. Varaiya · 1979
Earlier work this paper cites.
A new family of optimal adaptive controllers for Markov chains
P. R. Kumar and A. Becker · 1982
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1994
Earlier work this paper cites.
Constrained Markov decision processes
E. Altman · 1999
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcıa and F. Fernández · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Conservative bandits
Y. Wu, R. Shariff, T. Lattimore, and C. Szepesvári · 2016
Earlier work this paper cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
F. Berkenkamp, M. Turchetta, A. P. Schoellig, and A. Krause · 2017
Earlier work this paper cites.
Unifying pac and regret: uniform pac bounds for episodic reinforcement learning
C. Dann, T. Lattimore, and E. Brunskill · 2017
Earlier work this paper cites.
Conservative contextual linear bandits
A. Kazerouni, M. Ghavamzadeh, Y. Abbasi-Yadkori, and B. Van Roy · 2017
Cited alongside, same era.
Learning-based model predictive control for safe exploration
T. Koller, F. Berkenkamp, M. Turchetta, and A. Krause · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Q. Liang, F. Que, and E. Modiano · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Reward constrained policy optimization
C. Tessler, D. J. Mankowitz, and S. Mannor · 2018
Cited alongside, same era.
Safe exploration and optimization of constrained mdps using gaussian processes
A. Wachi, Y. Sui, Y. Yue, and M. Ono · 2018
Safe policies for reinforcement learning via primal-dual methods
S. Paternain, M. Calvo-Fullana, L. F. Chamon, and A. Ribeiro · 2019
Later among the works it cites.
Projection-based constrained policy optimization
T.-Y. Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge · 2019
Later among the works it cites.
Exploration-exploitation in constrained mdps
Y. Efroni, S. Mannor, and M. Pirotta · 2020
Later among the works it cites.
Improved algorithms for conservative exploration in bandits
E. Garcelon, M. Ghavamzadeh, A. Lazaric, and M. Pirotta · 2020
Later among the works it cites.
Learning adversarial markov decision processes with bandit feedback and unknown transition
C. Jin, T. Jin, H. Luo, S. Sra, and T. Yu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Linear stochastic bandits under safety constraints
S. Amani · 2019
Cited alongside, same era.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
R. Cheng, G. Orosz, R. M. Murray, and J. W. Burdick · 2019
Cited alongside, same era.
Challenges of real-world reinforcement learning
G. Dulac-Arnold, D. Mankowitz, and T. Hester · 2019
Cited alongside, same era.
Tight regret bounds for model-based reinforcement learning with greedy policies
Y. Efroni, N. Merlis, M. Ghavamzadeh, and S. Mannor · 2019
Cited alongside, same era.
Safe linear thompson sampling with side information
A. Moradipari, S. Amani, M. Alizadeh, and C. Thrampoulidis · 2019
Cited alongside, same era.
Upper confidence primal-dual reinforcement learning for cmdp with adversarial loss
S. Qiu, X. Wei, Z. Yang, J. Ye, and Z. Wang · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
A. Stooke, J. Achiam, and P. Abbeel · 2020
Later among the works it cites.
Constrained upper confidence reinforcement learning
L. Zheng and L. Ratliff · 2020
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
D. Ding, X. Wei, Z. Yang, Z. Wang, and M. Jovanovic · 2021
Closest in time.
An efficient pessimistic-optimistic algorithm for constrained linear bandits
X. Liu, B. Li, P. Shi, and L. Ying · 2021
Closest in time.
Stochastic bandits with linear constraints
A. Pacchiano, M. Ghavamzadeh, P. Bartlett, and H. Jiang · 2021
Closest in time.