Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has exceeded human performance in many synthetic settings such as video games and Go.
Markovian decision processes with uncertain transition probabilities
Jay K Satia and Roy E Lave Jr · 1973
Earlier work this paper cites.
Markov decision processes with imprecise transition probabilities
Chelsea C White III and Hany K Eldeib · 1994
Earlier work this paper cites.
Experts in a markov decision process
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2004
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Gergely Neu, Andras Antos, András György, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Lightning does not strike twice: robust mdps with coupled uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu · 2012
Earlier work this paper cites.
Theoretical statistics. lecture 12, 2013
Peter Bartlett · 2013
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Generalization and regularization in dqn
Jesse Farebrother, Marlos C Machado, and Michael Bowling · 2018
Earlier work this paper cites.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Song · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Cited alongside, same era.
Observational overfitting in reinforcement learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, and Behnam Neyshabur · 2019
rlberry - A Reinforcement Learning Library for Research and Education, 10 2021
Omar Darwiche Domingues, Yannis Flet-Berliac, Edouard Leurent, Pierre Ménard, Xuedong Shang, and Michal Valko · 2021
Later among the works it cites.
Corruption-robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Alex Slivkins, and Wen Sun · 2021
Later among the works it cites.
Decoupling value and policy for generalization in reinforcement learning
Roberta Raileanu and Rob Fergus · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Yue Wang and Shaofeng Zou · 2021
Later among the works it cites.
Wenhao Yang, Liangyu Zhang, and Zhihua Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Learning adversarial Markov decision processes with bandit feedback and unknown transition
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 2020
Cited alongside, same era.
Simultaneously learning stochastic and adversarial episodic mdps with known transition
Tiancheng Jin and Haipeng Luo · 2020
Cited alongside, same era.
Robust batch policy learning in markov decision processes
Zhengling Qi and Peng Liao · 2020
Cited alongside, same era.
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor · 2020
Cited alongside, same era.
Robust reinforcement learning using least squares policy iteration with provable performance guarantees
Kishan Panaganti Badrinath and Dileep Kalathil · 2021
Cited alongside, same era.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Zhengyuan Zhou, Qinxun Bai, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Later among the works it cites.
Doubly robust distributionally robust off-policy evaluation and learning
Nathan Kallus, Xiaojie Mao, Kaiwen Wang, and Zhengyuan Zhou · 2022
Closest in time.
Distributionally robust offline reinforcement learning with linear function approximation
Xiaoteng Ma, Zhipeng Liang, Li Xia, Jiheng Zhang, Jose Blanchet, Mingwen Liu, Qianchuan Zhao, and Zhengyuan Zhou · 2022
Closest in time.
Sample complexity of robust reinforcement learning with a generative model
Kishan Panaganti and Dileep Kalathil · 2022
Closest in time.
Policy gradient method for robust reinforcement learning
Yue Wang and Shaofeng Zou · 2022
Closest in time.
Nearly optimal policy optimization with stable at any time guarantee
Tianhao Wu, Yunchang Yang, Han Zhong, Liwei Wang, Simon Du, and Jiantao Jiao · 2022
Closest in time.
Corruption-robust offline reinforcement learning
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun · 2022
Closest in time.