Fetching the paper…
Reading the bibliography…
We adopt a policy optimization viewpoint towards policy evaluation for robust Markov decision process with $\mathrm{s}$-rectangular ambiguity sets.
Angenaherte auflosung von systemen linearer glei-chungen
Stefan Karczmarz · 1937
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J.N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
A convergent o (n) algorithm for off-policy temporal-difference learning with linear function approximation
Richard S Sutton, Csaba Szepesvári, and Hamid Reza Maei · 2008
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Yurii Nesterov · 2009
Earlier work this paper cites.
A randomized kaczmarz algorithm with exponential convergence
Thomas Strohmer and Roman Vershynin · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Risk-averse dynamic programming for markov decision processes
Andrzej Ruszczyński · 2010
Cited alongside, same era.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Cited alongside, same era.
Scaling up robust mdps using function approximation
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Cited alongside, same era.
Randomized iterative methods for linear systems
Robert M Gower and Peter Richtárik · 2015
Cited alongside, same era.
Random matrices: universality of local spectral statistics of non-hermitian matrices
Terence Tao and Van Vu · 2015
Cited alongside, same era.
Robust mdps with k-rectangular uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu · 2016
Cited alongside, same era.
Distributionally robust optimal control and mdp modeling
Alexander Shapiro · 2021
Later among the works it cites.
Policy optimization over general state and action spaces
Guanghui Lan · 2022
Later among the works it cites.
First-order policy optimization for robust markov decision process
Yan Li, Tuo Zhao, and Guanghui Lan · 2022
Later among the works it cites.
Distributionally robust q q -learning
Zijian Liu, Qinxun Bai, Jose Blanchet, Perry Dong, Wei Xu, Zhengqing Zhou, and Zhengyuan Zhou · 2022
Later among the works it cites.
Policy gradient method for robust reinforcement learning
Yue Wang and Shaofeng Zou · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning under model mismatch
Aurko Roy, Huan Xu, and Sebastian Pokutta · 2017
Cited alongside, same era.
Georgios Kotsalis, Guanghui Lan, and Tianjiao Li · 2020
Cited alongside, same era.
Robust markov decision processes: Beyond rectangularity
Vineet Goyal and Julien Grand-Clement · 2023
Closest in time.
Accelerated and instance-optimal policy evaluation with linear function approximation
Tianjiao Li, Guanghui Lan, and Ashwin Pananjady · 2023
Closest in time.
Policy mirror descent inherently explores action space
Yan Li and Guanghui Lan · 2023
Closest in time.