Fetching the paper…
Reading the bibliography…
A Budgeted Markov Decision Process (BMDP) is an extension of a Markov Decision Process to critical applications requiring safety constraints.
Optimal policies for controlled markov chains with a constraint
Frederick J. Beutler and Keith W. Ross · 1985
Earlier work this paper cites.
Constrained Markov Decision Processes
Eitan Altman · 1999
Earlier work this paper cites.
Beyond VaR: from measuring risk to managing risk
H. Mausser and D. Rosen · 2003
Earlier work this paper cites.
Tree-Based Batch Mode Reinforcement Learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Peter Geibel and Fritz Wysotzki · 2005
Earlier work this paper cites.
Robust Dynamic Programming
Garud N. Iyengar · 2005
Earlier work this paper cites.
Robust Control of Markov Decision Processes with Uncertain Transition Matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Neural fitted Q iteration - First experiences with a data efficient neural Reinforcement Learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Reinforcement learning for dialog management using least-squares policy iteration and fast feature selection
Lihong Li, Jason D. Williams, and Suhrid Balakrishnan · 2009
Earlier work this paper cites.
Optimizing debt collections using constrained reinforcement learning
Naoki Abe et al · 2010
Earlier work this paper cites.
Optimizing spoken dialogue management with fitted value iteration
Senthilkumar Chandramohan, Matthieu Geist, and Olivier Pietquin · 2010
Earlier work this paper cites.
Sample-efficient batch reinforcement learning for dialogue management optimization
Olivier Pietquin, Matthieu Geist, Senthilkumar Chandramohan, and Hervé Frezza-Buet · 2011
Cited alongside, same era.
Function approximation for continuous constrained mdps
Aditya Undurti, Alborz Geramifard, and Jonathan P. How · 2011
Cited alongside, same era.
Policy Gradients with Variance Related Risk Criteria
Aviv Tamar, Dotan Di Castro , and Shie Mannor · 2012
Cited alongside, same era.
Investment science
David G. Luenberger · 2013
Cited alongside, same era.
A survey of multi-objective sequential decision-making
Diederik M. Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Cited alongside, same era.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Approximate linear programming for constrained partially observable markov decision processes
Pascal Poupart, Aarti Malhotra, Pei Pei, Kee-Eung Kim, Bongseok Goh, and Michael Bowling · 2015
Later among the works it cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Later among the works it cites.
Budget allocation using weakly coupled, constrained markov decision processes
Craig Boutilier and Tyler Lu · 2016
Later among the works it cites.
Safe policy improvement by minimizing robust baseline regret
Mohammad Petrik, Marek Ghavamzadeh, , and Yinlam Chow · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Later among the works it cites.
Safe transfer learning for dialogue applications
Nicolas Carrara, Romain Laroche, Jean-Léon Bouraoui, Tanguy Urvoy, and Olivier Pietquin · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multiobjective Reinforcement Learning: A Comprehensive Overview
Chunming Liu, Xin Xu, and Dewen Hu · 2014
Cited alongside, same era.
Risk-Sensitive and Robust Decision-Making: a CVaR Optimization Approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Cited alongside, same era.
A Comprehensive Survey on Safe Reinforcement Learning
Javier García and Fernando Fernández · 2015
Cited alongside, same era.
Optimising turn-taking strategies with reinforcement learning.
Hatim Khouzaimi, Romain Laroche, and Fabrice. Lefevre · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Later among the works it cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2018
Later among the works it cites.
Approximate Robust Control of Uncertain Dynamical Systems
Edouard Leurent, Yann Blanco, Denis Efimov, and Odalric-Ambrym Maillard · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Closest in time.
Safe policy improvement with baseline bootstrapping
Romain Laroche and Rémi Trichelair, Paul and Tachet des Combes · 2019
Closest in time.
Batch policy learning under constraints
Hoang M. Le, Cameron Voloshin, and Yisong Yue · 2019
Closest in time.