Fetching the paper…
Reading the bibliography…
In this paper, we study the non-asymptotic and asymptotic performances of the optimal robust policy and value function of robust Markov Decision Processes(MDPs), where the optimal robust policy and value function are solved only from a generative model.
Neuro-dynamic programming: an overview
Dimitri P Bertsekas and John N Tsitsiklis · 1995
Earlier work this paper cites.
Weak convergence
Aad W Van Der Vaart and Jon A Wellner · 1996
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
Recursive multiple-priors
Larry G Epstein and Martin Schneider · 2003
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Keisuke Hirano, Guido W Imbens, and Geert Ridder · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Bias and variance in value function estimation
Shie Mannor, Duncan Simester, Peng Sun, and John N Tsitsiklis · 2004
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
The robustness-performance tradeoff in markov decision processes
Huan Xu and Shie Mannor · 2006
Earlier work this paper cites.
Model-based function approximation in reinforcement learning
Nicholas K Jong and Peter Stone · 2007
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman · 2009
Earlier work this paper cites.
Parametric regret in uncertain markov decision processes
Huan Xu and Shie Mannor · 2009
Earlier work this paper cites.
Distributionally robust optimization under moment uncertainty with application to data-driven problems
Erick Delage and Yinyu Ye · 2010
Earlier work this paper cites.
Modelling transition dynamics in mdps with rkhs embeddings
Steffen Grünewälder, Guy Lever, Luca Baldassarre, Massimilano Pontil, and Arthur Gretton · 2012
Earlier work this paper cites.
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Earlier work this paper cites.
Lightning does not strike twice: robust mdps with coupled uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu · 2012
Earlier work this paper cites.
Approximate dynamic programming by minimizing distributionally robust bounds
Marek Petrik · 2012
Earlier work this paper cites.
A framework for optimization under ambiguity
David Wozabal · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Robust solutions of optimization problems affected by uncertain probabilities
Aharon Ben-Tal, Dick Den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen · 2013
Earlier work this paper cites.
Robust modified policy iteration
David L Kaufman and Andrew J Schaefer · 2013
Earlier work this paper cites.
Reinforcement learning in robust markov decision processes
Shiau Hong Lim, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Ber c · 2013
Earlier work this paper cites.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, Jan Peters, et al · 2014
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, Lihong Li, et al · 2014
Earlier work this paper cites.
Raam: The benefits of robustness in approximating aggregated mdps in reinforcement learning
Marek Petrik and Dharmashankar Subramanian · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
The self-normalized estimator for counterfactual learning
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Later among the works it cites.
Beyond confidence regions: Tight bayesian ambiguity sets for robust mdps
Marek Petrik and Reazul Hasan Russel · 2019
Later among the works it cites.
Distributionally robust reinforcement learning
Elena Smirnova, Elvis Dohmatob, and Jérémie Mary · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Later among the works it cites.
Coindice: Off-policy confidence interval estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Distributionally robust stochastic optimization with wasserstein distance
Rui Gao and Anton J Kleywegt · 2016
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
Mohammad Ghavamzadeh, Marek Petrik, and Yinlam Chow · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Robust sensitivity analysis for stochastic systems
Henry Lam · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
f f -divergence inequalities
Igal Sason and Sergio Verdú · 2016
Cited alongside, same era.
Bo Dai, Ofir Nachum, Yinlam Chow, Lihong Li, Csaba Szepesvári, and Dale Schuurmans · 2020
Later among the works it cites.
Distributional robustness and regularization in reinforcement learning
Esther Derman and Shie Mannor · 2020
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Nathan Kallus and Masatoshi Uehara · 2020
Later among the works it cites.
Distributionally robust policy evaluation and learning in offline contextual bandits
Nian Si, Fan Zhang, Zhengyuan Zhou, and Jose Blanchet · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean Foster, and Sham M Kakade · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Later among the works it cites.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund · 2020
Later among the works it cites.
Optimizing percentile criterion using robust mdps
Bahram Behzadian, Reazul Hasan Russel, Marek Petrik, and Chin Pang Ho · 2021
Closest in time.
Risk bounds and rademacher complexity in batch reinforcement learning
Yaqi Duan, Chi Jin, and Zhiyuan Li · 2021
Closest in time.
Learning models with uniform performance via distributionally robust optimization
John C Duchi and Hongseok Namkoong · 2021
Closest in time.
Statistics of robust optimization: A generalized empirical likelihood approach
John C Duchi, Peter W Glynn, and Hongseok Namkoong · 2021
Closest in time.
Partial policy iteration for l1-robust markov decision processes
Chin Pang Ho, Marek Petrik, and Wolfram Wiesemann · 2021
Closest in time.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Closest in time.
Polyak-ruppert averaged q-leaning is statistically efficient
Xiang Li, Wenhao Yang, Zhihua Zhang, and Michael I Jordan · 2021
Closest in time.
On the optimality of batch policy optimization algorithms
Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvari, and Dale Schuurmans · 2021
Closest in time.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Closest in time.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Closest in time.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Qinxun Bai, Zhengyuan Zhou, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Closest in time.
Robust markov decision processes: Beyond rectangularity
Vineet Goyal and Julien Grand-Clement · 2022
Closest in time.
Sample complexity of robust reinforcement learning with a generative model
Kishan Panaganti and Dileep Kalathil · 2022
Closest in time.