Fetching the paper…
Reading the bibliography…
In this paper, we focus on the problem of robustifying reinforcement learning (RL) algorithms with respect to model uncertainties.
Reinforcement learning
Sutton, R. S. and Barto, A · 1998
Earlier work this paper cites.
Foundations of Inventory Management
Zipkin, P. H · 2000
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P · 2003
Earlier work this paper cites.
Constrained Markov Decision Processes
Altman, E · 2004
Earlier work this paper cites.
Robust solutions to Markov decision problems with uncertain transition matrices
Nilim, A. and Ghaoui, L. E · 2004
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Geibel, P. and Wysotzki, F · 2005
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 2005
Earlier work this paper cites.
Robust, Risk-Sensitive, and Data-driven Control of Markov Decision Processes
Le Tallec, Y · 2007
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint
Borkar, V. S · 2009
Cited alongside, same era.
Approximate dynamic programming by minimizing distributionally robust bounds
Petrik, M · 2012
Cited alongside, same era.
Robust Markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Cited alongside, same era.
Autonomy and machine intelligence in complex systems: A tutorial
Vamvoudakis, K., Antsaklis, P., Dixon, W., Hespanha, J., Lewis, F., Modares, H., and Kiumarsi, B · 2015
Cited alongside, same era.
Convex synthesis of randomized policies for controlled markov chains with density safety upper bound constraints
Chamiea, M. E., Yu, Y., and Acikmese, B · 2016
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
Model‐based vs data‐driven adaptive control: An overview
Benosman, M · 2018
Later among the works it cites.
Safe exploration in continuous action spaces, 2018
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Later among the works it cites.
High-Confidence Policy Optimization: Reshaping Ambiguity Sets in Robust MDPs
Behzadian, B., Russel, R. H., and Petrik, M · 2019
Later among the works it cites.
Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs
Petrik, M. and Russell, R. H · 2019
Later among the works it cites.
Beyond confidence regions: Tight Bayesian ambiguity sets for robust MDPs
Russel, R. H. and Petrik, M · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Interpretable Policies for Dynamic Product Recommendations
Petrik, M. and Luss, R · 2016
Cited alongside, same era.
One-shot visual imitation learning via meta-learning
Finn, C., Yu, T., Zhang, T., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Later among the works it cites.
Sim-to-real transfer learning using robustified controllers in robotic tasks involving complex dynamics
van Baar, J., Sullivan, A., Corcodel, R., Jha, D., Romeres, D., and Nikovski, D. N · 2019
Later among the works it cites.