Fetching the paper…
Reading the bibliography…
Robust Markov Decision Processes (RMDPs) are a widely used framework for sequential decision-making under parameter uncertainty.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Recursive games
Hugh Everett · 1957
Earlier work this paper cites.
Stochastic games with zero stop probabilities
Dean Gillette · 1957
Earlier work this paper cites.
The big match
David Blackwell and Tom S Ferguson · 1968
Earlier work this paper cites.
Stochastic games with perfect information and time average payoff
Thomas M Liggett and Steven A Lippman · 1969
Earlier work this paper cites.
Markov decision processes with uncertain transition probabilities
J.K. Satia and R.L. Lave · 1973
Earlier work this paper cites.
Repeated games with absorbing states
Elon Kohlberg · 1974
Earlier work this paper cites.
The asymptotic theory of stochastic games
Truman Bewley and Elon Kohlberg · 1976
Earlier work this paper cites.
An expected average reward criterion
K-J Bierth · 1987
Earlier work this paper cites.
The complexity of stochastic games
Anne Condon · 1992
Earlier work this paper cites.
On the generation of markov decision processes
TW Archibald, KIM McKinnon, and LC Thomas · 1995
Earlier work this paper cites.
Geometric categories and o-minimal structures
Lou Van den Dries and Chris Miller · 1996
Earlier work this paper cites.
The complexity of mean payoff games on graphs
Uri Zwick and Mike Paterson · 1996
Earlier work this paper cites.
Bounded parameter Markov decision processes
Robert Givan, Sonia Leach, and Thomas Dean · 1997
Earlier work this paper cites.
Application of stochastic dynamic programming to optimal fire management of a spatially structured threatened species
Hugh Possingham and G Tuck · 1997
Earlier work this paper cites.
O-minimal structures and real analytic geometry
Lou Van Den Dries · 1998
Earlier work this paper cites.
An introduction to o-minimal geometry
Michel Coste · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
A first course on zero-sum repeated games
Sylvain Sorin · 2002
Earlier work this paper cites.
Spectral theorem for convex monotone homogeneous maps, and ergodic control
Marianne Akian and Stéphane Gaubert · 2003
Earlier work this paper cites.
Isometric logratio transformations for compositional data analysis
Juan José Egozcue, Vera Pawlowsky-Glahn, Glòria Mateu-Figueras, and Carles Barcelo-Vidal · 2003
Earlier work this paper cites.
An algorithm to identify and compute average optimal policies in multichain Markov decision processes
Arie Leizarowitz · 2003
Earlier work this paper cites.
Stochastic games and applications
Abraham Neyman, Sylvain Sorin, and S Sorin · 2003
Earlier work this paper cites.
Robust dynamic programming
G. Iyengar · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition probabilities
A. Nilim and L. El Ghaoui · 2005
Earlier work this paper cites.
Robust, risk-sensitive, and data-driven control of Markov decision processes
Yann Le Tallec · 2007
Cited alongside, same era.
Bounded parameter Markov decision processes with average reward criterion
Ambuj Tewari and Peter L Bartlett · 2007
Cited alongside, same era.
The complexity of solving stochastic games on graphs
Daniel Andersson and Peter Bro Miltersen · 2009
Cited alongside, same era.
Percentile optimization for Markov decision processes with parameter uncertainty
Erick Delage and Shie Mannor · 2010
Cited alongside, same era.
Distributionally robust Markov decision processes
Huan Xu and Shie Mannor · 2010
Cited alongside, same era.
Markov decision processes with applications to finance
Nicole Bäuerle and Ulrich Rieder · 2011
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
The operator approach to entropy games
Marianne Akian, Stéphane Gaubert, Julien Grand-Clément, and Jérémie Guillaud · 2019
Later among the works it cites.
A tutorial on zero-sum stochastic games
Jérôme Renault · 2019
Later among the works it cites.
Average-reward model-free reinforcement learning: a systematic review and literature mapping
Vektor Dewanto, George Dunn, Ali Eshragh, Marcus Gallagher, and Fred Roosta · 2020
Later among the works it cites.
Fast algorithms for l ∞ l_{\infty} constrained s-rectangular robust MDPs
Bahram Behzadian, Marek Petrik, and Chin Pang Ho · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Handbook of Markov decision processes: methods and applications
Eugene A Feinberg and Adam Shwartz · 2012
Cited alongside, same era.
Competitive Markov decision processes
Jerzy Filar and Koos Vrieze · 2012
Cited alongside, same era.
Artificial intelligence framework for simulating clinical decision-making: A Markov decision process approach
Casey C Bennett and Kris Hauser · 2013
Cited alongside, same era.
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
Thomas Dueholm Hansen, Peter Bro Miltersen, and Uri Zwick · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Robust Markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Cited alongside, same era.
Julien Grand-Clément and Christian Kroer · 2021
Later among the works it cites.
First-order methods for wasserstein distributionally robust MDP
Julien Grand-Clement and Christian Kroer · 2021
Later among the works it cites.
Partial policy iteration for l1-robust Markov decision processes
Chin Pang Ho, Marek Petrik, and Wolfram Wiesemann · 2021
Later among the works it cites.
Robust inverse reinforcement learning under transition dynamics mismatch
Luca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Adrian Weller, and Volkan Cevher · 2021
Later among the works it cites.
Topics in optimal transportation
Cédric Villani · 2021
Later among the works it cites.
Robust imitation learning against variations in environment dynamics
Jongseong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho, Sungho Choi, and Youngchul Sung · 2022
Later among the works it cites.
Robust Markov decision processes: Beyond rectangularity
Vineet Goyal and Julien Grand-Clément · 2022
Later among the works it cites.
Robust phi-divergence mdps
Chin Pang Ho, Marek Petrik, and Wolfram Wiesemann · 2022
Later among the works it cites.
First-order policy optimization for robust Markov decision process
Yan Li, Tuo Zhao, and Guanghui Lan · 2022
Later among the works it cites.
Sample complexity of robust reinforcement learning with a generative model
Kishan Panaganti and Dileep Kalathil · 2022
Later among the works it cites.
Robust markov decision processes with data-driven, distance-based ambiguity sets
Sivaramakrishnan Ramani and Archis Ghate · 2022
Later among the works it cites.
Robustness of proactive intensive care unit transfer policies
Julien Grand-Clément, Carri W Chan, Vineet Goyal, and Gabriel Escobar · 2023
Closest in time.
Solving optimization problems with blackwell approachability
Julien Grand-Clément and Christian Kroer · 2023
Closest in time.
Policy gradient for rectangular robust markov decision processes
Navdeep Kumar, Esther Derman, Matthieu Geist, Kfir Y Levy, and Shie Mannor · 2023
Closest in time.
Policy gradient algorithms for robust MDPs with non-rectangular uncertainty sets
Mengmeng Li, Tobias Sutter, and Daniel Kuhn · 2023
Closest in time.
Policy gradient in robust mdps with global convergence guarantee
Qiuhao Wang, Chin Pang Ho, and Marek Petrik · 2023
Closest in time.
Robust average-reward markov decision processes
Yue Wang, Alvaro Velasquez, George Atia, Ashley Prater-Bennette, and Shaofeng Zou · 2023
Closest in time.
On the convex formulations of robust Markov decision processes
Julien Grand-Clément and Marek Petrik · 2024
Closest in time.
Reducing blackwell and average optimality to discounted mdps via the blackwell discount factor
Julien Grand-Clément and Marek Petrik · 2024
Closest in time.
A family of-rectangular robust mdps: Relative conservativeness, asymptotic analyses, and finite-sample properties
Sivaramakrishnan Ramani and Archis Ghate · 2024
Closest in time.