Fetching the paper…
Reading the bibliography…
Constrained Markov Decision Processes are a class of stochastic decision problems in which the decision maker must select a policy that satisfies auxiliary cost constraints.
Linear and nonlinear programming , volume 2
David G Luenberger, Yinyu Ye, et al · 1984
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state markov decision processes
Stefano P Coraluppi and Steven I Marcus · 1999
Earlier work this paper cites.
An mdp-based recommender system
Guy Shani, David Heckerman, and Ronen I Brafman · 2005
Earlier work this paper cites.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Chernoff-hoeffding bounds for markov chains: Generalized and simplified
Kai-Min Chung, Henry Lam, Zhenming Liu, and Michael Mitzenmacher · 2012
Earlier work this paper cites.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Multi-armed bandit with budget constraint and variable costs
Wenkui Ding, Tao Qin, Xu-Dong Zhang, and Tie-Yan Liu · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Later among the works it cites.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Later among the works it cites.
Deep reinforcement learning framework for autonomous driving
Ahmad EL Sallab, Mohammed Abdou, Etienne Perot, and Senthil Yogamani · 2017
Later among the works it cites.
Learning-based model predictive control for safe exploration
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause · 2018
Later among the works it cites.
Safe exploration and optimization of constrained mdps using gaussian processes
Akifumi Wachi, Yanan Sui, Yisong Yue, and Masahiro Ono · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Fairness in learning: Classic and contextual bandits
Matthew Joseph, Michael Kearns, Jamie H Morgenstern, and Aaron Roth · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone
Cited in the paper.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh
Cited in the paper.
Budget-constrained multi-armed bandits with multiple plays
Datong P Zhou and Claire J Tomlin · 2018
Later among the works it cites.
Richard Cheng, Gábor Orosz, Richard M Murray, and Joel W Burdick · 2019
Later among the works it cites.
Controlled markov processes with safety state constraints
Mahmoud El Chamie, Yue Yu, Behçet Açıkmeşe, and Masahiro Ono · 2019
Later among the works it cites.