Fetching the paper…
Reading the bibliography…
Motivated by the need for a robust policy in the face of environment shifts between training and deployment, we contribute to the theoretical foundation of distributionally robust reinforcement learning (DRRL).
Technical note—time inconsistency of optimal policies of distributionally robust inventory models
Alexander Shapiro and Linwei Xin · 1932
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
The theory of dynamic programming
Richard Bellman · 1954
Earlier work this paper cites.
On general minimax theorems
Maurice Sion · 1958
Earlier work this paper cites.
Equilibrium in a stochastic n n -person game
Arlington M Fink · 1964
Earlier work this paper cites.
Equilibrium points of stochastic non-cooperative n n -person games
Masayuki Takahashi · 1964
Earlier work this paper cites.
Existence of stationary equilibrium strategies in non-zero sum discounted stochastic games with uncountable state space and state-independent transitions
T Parthasarathy and S Sinha · 1989
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Markov perfect equilibrium: I. observable actions
Eric Maskin and Jean Tirole · 2001
Earlier work this paper cites.
Minimax control of discrete-time stochastic systems
JI González-Trejo, Onésimo Hernández-Lerma, and Luis F Hoyos-Reyes · 2002
Earlier work this paper cites.
On a new class of nonzero-sum discounted stochastic games having stationary nash equilibrium points
Andrzej S Nowak · 2003
Earlier work this paper cites.
Games of incomplete information, ergodic theory, and the measurability of equilibria
Robert Samuel Simon · 2003
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Dynamic pricing: A learning approach
Dimitris Bertsimas and Georgia Perakis · 2006
Earlier work this paper cites.
Robust, risk-sensitive, and data-driven control of Markov decision processes
Yann Le Tallec · 2007
Earlier work this paper cites.
Distributionally robust markov decision processes
Huan Xu and Shie Mannor · 2010
Earlier work this paper cites.
Reinforcement learning in robotics: Applications and real-world challenges
Petar Kormushev, Sylvain Calinon, and Darwin G Caldwell · 2013
Earlier work this paper cites.
Discounted stochastic games with no stationary nash equilibrium: two examples
Yehuda Levy · 2013
Earlier work this paper cites.
Common information based markov perfect equilibria for stochastic games with asymmetric information: Finite games
Ashutosh Nayyar, Abhishek Gupta, Cedric Langbort, and Tamer Başar · 2013
Earlier work this paper cites.
The simplex method is strongly polynomial for deterministic markov decision processes, 2013
Ian Post and Yinyu Ye · 2013
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Cited alongside, same era.
Dynamic treatment regimes
Bibhas Chakraborty and Susan A Murphy · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Stochastic games
Eilon Solan and Nicolas Vieille · 2015
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Real-time bidding by reinforcement learning in display advertising
Han Cai, Kan Ren, Weinan Zhang, Kleanthis Malialis, Jun Wang, Yong Yu, and Defeng Guo · 2017
Lectures on stochastic programming: modeling and theory
Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczynski · 2021
Later among the works it cites.
Wenhao Yang, Liangyu Zhang, and Zhihua Zhang · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Zhengyuan Zhou, Qinxun Bai, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2022
Later among the works it cites.
Robust markov decision processes: Beyond rectangularity
Vineet Goyal and Julien Grand-Clément · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A study of distributionally robust multistage stochastic optimization
Jianqiu Huang, Kezhuo Zhou, and Yongpei Guan · 2017
Cited alongside, same era.
Information structures and values in zero-sum stochastic games
Ashutosh Nayyar and Abhishek Gupta · 2017
Cited alongside, same era.
Virtual to real reinforcement learning for autonomous driving
Xinlei Pan, Yurong You, Ziyan Wang, and Cewu Lu · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Deep reinforcement learning for sponsored search real-time bidding
Jun Zhao, Guang Qiu, Ziyu Guan, Wei Zhao, and Xiaofei He · 2018
Cited alongside, same era.
Distributionally robust Q-learning
Zijian Liu, Qinxun Bai, Jose Blanchet, Perry Dong, Wei Xu, Zhengqing Zhou, and Zhengyuan Zhou · 2022
Later among the works it cites.
Distributionally robust modeling of optimal control
Alexander Shapiro · 2022
Later among the works it cites.
Laixi Shi and Yuejie Chi · 2022
Later among the works it cites.
Double pessimism is provably efficient for distributionally robust offline reinforcement learning: Generic algorithm and robust partial coverage, 2023
Jose Blanchet, Miao Lu, Tong Zhang, and Han Zhong · 2023
Closest in time.
Seeing is not believing: Robust reinforcement learning against spurious correlation
Wenhao Ding, Laixi Shi, Yuejie Chi, and Ding Zhao · 2023
Closest in time.
Bandits atop reinforcement learning: Tackling online inventory models with cyclic demands
Xiao-Yue Gong and David Simchi-Levi · 2023
Closest in time.
Beyond discounted returns: Robust markov decision processes with average and blackwell optimality
Julien Grand-Clement, Marek Petrik, and Nicolas Vieille · 2023
Closest in time.
Rectangularity and duality of distributionally robust markov decision processes, 2023
Yan Li and Alexander Shapiro · 2023
Closest in time.
Wasserstein distributionally robust linear-quadratic estimation under martingale constraints
Kyriakos Lotidis, Nicholas Bambos, Jose Blanchet, and Jiajin Li · 2023
Closest in time.
Ambiguous dynamic treatment regimes: A reinforcement learning approach
Soroush Saghafian · 2023
Closest in time.
The curious price of distributional robustness in reinforcement learning with a generative model, 2023
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, Matthieu Geist, and Yuejie Chi · 2023
Closest in time.
Distributionally robust linear quadratic control, 2023
Bahar Taşkesen, Dan A. Iancu, Çağıl Koçyiğit, and Daniel Kuhn · 2023
Closest in time.
Improved sample complexity bounds for distributionally robust reinforcement learning, 2023
Zaiyan Xu, Kishan Panaganti, and Dileep Kalathil · 2023
Closest in time.
Avoiding model estimation in robust markov decision processes with a generative model, 2023
Wenhao Yang, Han Wang, Tadashi Kozuno, Scott M. Jordan, and Zhihua Zhang · 2023
Closest in time.
Data-driven hospital admission control: A learning approach
Mohammad Zhalechian, Esmaeil Keyvanshokooh, Cong Shi, and Mark P Van Oyen · 2023
Closest in time.
A reinforcement learning approach to view planning for automated inspection tasks
Christian Landgraf, Bernd Meese, Michael Pabst, Georg Martius, and Marco F. Huber · 2030
Closest in time.