Fetching the paper…
Reading the bibliography…
We consider the problem of solving robust Markov decision process (MDP), which involves a set of discounted, finite state, finite action space MDPs with uncertain transition kernels.
The theory of max-min and its application to weapons allocation problems
John M. Danskin · 1967
Earlier work this paper cites.
Convex analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Strong convexity of sets and functions
Jean-Philippe Vial · 1982
Earlier work this paper cites.
Lectures on Geometric Measure Theory
Leon Simon · 1983
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
On fréchet subdifferentials
A Ya Kruger · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2003
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the curses of dimensionality
Warren B Powell · 2007
Earlier work this paper cites.
Robust regression and lasso
Huan Xu, Constantine Caramanis, and Shie Mannor · 2008
Earlier work this paper cites.
Robustness and regularization of support vector machines
Huan Xu, Constantine Caramanis, and Shie Mannor · 2009
Earlier work this paper cites.
Risk-averse dynamic programming for markov decision processes
Andrzej Ruszczyński · 2010
Earlier work this paper cites.
Robust modified policy iteration
David L Kaufman and Andrew J Schaefer · 2013
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Introduction to geometric measure theory
Leon Simon · 2014
Cited alongside, same era.
Scaling up robust mdps using function approximation
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Cited alongside, same era.
Markov chains and mixing times
David A Levin and Yuval Peres · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Reinforcement learning under model mismatch
Aurko Roy, Huan Xu, and Sebastian Pokutta · 2017
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Twice regularized mdps and the equivalence between robustness and regularization
Esther Derman, Matthieu Geist, and Shie Mannor · 2021
Later among the works it cites.
Partial policy iteration for l1-robust markov decision processes
Chin Pang Ho, Marek Petrik, and Wolfram Wiesemann · 2021
Later among the works it cites.
On the linear convergence of natural policy gradient algorithm
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2021
Later among the works it cites.
Guanghui Lan · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Yue Wang and Shaofeng Zou · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Cited alongside, same era.
Scalable first-order methods for robust mdps
Julien Grand-Clément and Christian Kroer · 2020
Cited alongside, same era.
Risk-averse learning by temporal difference methods
Umit Kose and Andrzej Ruszczynski · 2020
Cited alongside, same era.
First-order and stochastic optimization methods for machine learning
Guanghui Lan · 2020
Cited alongside, same era.
Implicit bias of gradient descent based adversarial training on separable data
Yan Li, Ethan X.Fang, Huan Xu, and Tuo Zhao · 2020
Cited alongside, same era.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Cited alongside, same era.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Zhengyuan Zhou, Qinxun Bai, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Later among the works it cites.
A first-order approach to accelerated value iteration
Vineet Goyal and Julien Grand-Clement · 2022
Closest in time.
Robust markov decision processes: Beyond rectangularity
Vineet Goyal and Julien Grand-Clement · 2022
Closest in time.
Efficient policy iteration for robust markov decision processes via regularization
Navdeep Kumar, Kfir Levy, Kaixin Wang, and Shie Mannor · 2022
Closest in time.
Yan Li, Tuo Zhao, and Guanghui Lan · 2022
Closest in time.
Distributionally robust q q -learning
Zijian Liu, Qinxun Bai, Jose Blanchet, Perry Dong, Wei Xu, Zhengqing Zhou, and Zhengyuan Zhou · 2022
Closest in time.
Distributionally robust offline reinforcement learning with linear function approximation
Xiaoteng Ma, Zhipeng Liang, Li Xia, Jiheng Zhang, Jose Blanchet, Mingwen Liu, Qianchuan Zhao, and Zhengyuan Zhou · 2022
Closest in time.
Sample complexity of robust reinforcement learning with a generative model
Kishan Panaganti and Dileep Kalathil · 2022
Closest in time.
Policy gradient method for robust reinforcement learning
Yue Wang and Shaofeng Zou · 2022
Closest in time.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Closest in time.
Rorl: Robust offline reinforcement learning via conservative smoothing
Rui Yang, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang, Chongjie Zhang, and Lei Han · 2022
Closest in time.
Policy mirror descent inherently explores action space
Yan Li and Guanghui Lan · 2023
Closest in time.