Fetching the paper…
Reading the bibliography…
Safety in reinforcement learning has become increasingly important in recent years.
Provably efficient q-learning with low switching cost
Bai, Y., Xie, T., Jiang, N., and Wang, Y.-X. (2019) · 1905
Earlier work this paper cites.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L. F., Calvo-Fullana, M., and Ribeiro, A. (2019b) · 1910
Earlier work this paper cites.
Moradipari, A., Amani, S., Alizadeh, M., and Thrampoulidis, C. (2019) · 1911
Earlier work this paper cites.
Safe policies for reinforcement learning via primal-dual methods
Paternain, S., Calvo-Fullana, M., Chamon, L. F., and Ribeiro, A. (2019a) · 1911
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Dynamic programming and optimal control: Vol. 1
Bertsekas, D. P. et al. (2000) · 2000
Earlier work this paper cites.
Constrained upper confidence reinforcement learning
Zheng, L. and Ratliff, L. J. (2020) · 2001
Earlier work this paper cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanović, M. R. (2020a) · 2003
Earlier work this paper cites.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Earlier work this paper cites.
Qiu, S., Wei, X., Yang, Z., Ye, J., and Wang, Z. (2020) · 2003
Earlier work this paper cites.
Stochastic bandits with linear constraints
Pacchiano, A., Ghavamzadeh, M., Bartlett, P., and Jiang, H. (2020) · 2006
Earlier work this paper cites.
Safe reinforcement learning via curriculum induction
Turchetta, M., Kolobov, A., Shah, S., Krause, A., and Agarwal, A. (2020) · 2006
Earlier work this paper cites.
A sample-efficient algorithm for episodic finite-horizon mdp with constraints
Kalagarla, K. C., Jain, R., and Nuzzo, P. (2020) · 2009
Cited alongside, same era.
Optimizing debt collections using constrained reinforcement learning
Abe, N., Melville, P., Pendus, C., Reddy, C. K., Jensen, D. L., Thomas, V. P., Bennett, J. J., Anderson, G. F., Cooley, B. R., Kowalczyk, M., et al. (2010) · 2010
Cited alongside, same era.
Algorithms for reinforcement learning
Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J. (2020) · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S. (2018) · 2018
Later among the works it cites.
Safe exploration and optimization of constrained mdps using gaussian processes
Wachi, A., Sui, Y., Yue, Y., and Ono, M. (2018) · 2018
Later among the works it cites.
Linear stochastic bandits under safety constraints
Amani, S., Alizadeh, M., and Thrampoulidis, C. (2019) · 2019
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M. (2019) · 2019
Later among the works it cites.
Convergent policy optimization for safe reinforcement learning
Yu, M., Yang, Z., Kolar, M., and Wang, Z. (2019) · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xu, T., Liang, Y., and Lan, G. (2020) · 2011
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Safe exploration in finite markov decision processes with gaussian processes
Turchetta, M., Berkenkamp, F., and Krause, A. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A. (2017) · 2017
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M. (2018) · 2018
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M. (2020b)
Cited in the paper.
Later among the works it cites.
Conservative exploration in reinforcement learning
Garcelon, E., Ghavamzadeh, M., Lazaric, A., and Pirotta, M. (2020) · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2020) · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P. (2020) · 2020
Later among the works it cites.
Safe reinforcement learning in constrained markov decision processes
Wachi, A. and Sui, Y. (2020) · 2020
Later among the works it cites.
Decentralized multi-agent linear bandits with safety constraints
Amani, S. and Thrampoulidis, C. (2021) · 2021
Closest in time.