Fetching the paper…
Reading the bibliography…
Safety and robustness are two desired properties for any reinforcement learning algorithm.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Lyapunov-Constrained Action Sets for Reinforcement Learning
Perkins, T. J. and Barto, A. G · 2000
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P · 2003
Earlier work this paper cites.
Active Fault Tolerant Control Systems: Stochastic Analysis and Synthesis
Mahmoud, M. M., Jiang, J., and Zhang, Y · 2003
Earlier work this paper cites.
Constrained Markov Decision Processes
Altman, E · 2004
Earlier work this paper cites.
Robust solutions to Markov decision problems with uncertain transition matrices
Nilim, A. and Ghaoui, L. E · 2004
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Geibel, P. and Wysotzki, F · 2005
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Nilim, A. and Ghaoui, L. E · 2005
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2005
Earlier work this paper cites.
Robust, Risk-Sensitive, and Data-driven Control of Markov Decision Processes
Le Tallec, Y · 2007
Earlier work this paper cites.
Nonlinear dynamical systems and control: A Lyapunov-based approach
Haddad, W. M · 2008
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint
Borkar, V. S · 2009
Cited alongside, same era.
Algorithms for Reinforcement Learning
Szepesvári, C · 2010
Cited alongside, same era.
Robust Markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Cited alongside, same era.
Algorithms for CVaR optimization in MDPs
Chow, Y. and Ghavamzadeh, M · 2014
Cited alongside, same era.
Optimizing the CVaR via Sampling
Tamar, A., Glassner, Y., and Mannor, S · 2014
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
Model‐based vs data‐driven adaptive control: An overview
Benosman, M · 2018
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Later among the works it cites.
Safe exploration in continuous action spaces, 2018
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Later among the works it cites.
Soft-robust actor-critic policy-gradient
Derman, E., Mankowitz, D. J., Mann, T. A., and Mannor, S · 2018
Later among the works it cites.
The architectural implications of autonomous driving: Constraints and acceleration
Lin, S. C., Zhang, Y., Hsu, C. H., Skach, M., Haque, M. E., Tang, L., and Mars, J · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Autonomy and machine intelligence in complex systems: A tutorial
Vamvoudakis, K., Antsaklis, P., Dixon, W., Hespanha, J., Lewis, F., Modares, H., and Kiumarsi, B · 2015
Cited alongside, same era.
Convex synthesis of randomized policies for controlled markov chains with density safety upper bound constraints
Chamiea, M. E., Yu, Y., and Acikmese, B · 2016
Cited alongside, same era.
Constrained Policy Optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
Towards stability in learning based control: A bayesian optimization based adaptive controller
Farahmand, A.-M. and Benosman, M · 2017
Cited alongside, same era.
In Levine, S., Vanhoucke, V., and Goldberg, K. (eds.), One-Shot Visual Imitation Learning via Meta-Learning , 2017
Finn, C., Yu, T., Zhang, T., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs
Russel, R. H. and Petrik, M
Cited in the paper.
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L. F., Calvo-Fullana, M., and Ribeiro, A · 2019
Later among the works it cites.
Sim-to-real transfer learning using robustified controllers in robotic tasks involving complex dynamics
van Baar, J., Sullivan, A., Corcodel, R., Jha, D., Romeres, D., and Nikovski, D. N · 2019
Later among the works it cites.
Optimizing Percentile Criterion Using Robust MDPs
Behzadian, B., Russel, R. H., Petrik, M., and Ho, C. P · 2021
Closest in time.