Fetching the paper…
Reading the bibliography…
Robust Markov Decision Processes (MDPs) are a powerful framework for modeling sequential decision-making problems with model uncertainty.
Problem complexity and method efficiency in optimization
Arkadi Nemirovski and David Yudin · 1983
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O(1/kˆ 2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Markov Decision Processes : Discrete Stochastic Dynamic Programming
M.L. Puterman · 1994
Earlier work this paper cites.
On the generation of Markov decision processes
TW Archibald, KIM McKinnon, and LC Thomas · 1995
Earlier work this paper cites.
Bounded parameter Markov decision processes
Robert Givan, Sonia Leach, and Thomas Dean · 1997
Earlier work this paper cites.
Applications of second-order cone programming
Miguel Sousa Lobo, Lieven Vandenberghe, Stephen Boyd, and Hervé Lebret · 1998
Earlier work this paper cites.
Robust solutions of linear programming problems contaminated with uncertain data
Aharon Ben-Tal and Arkadi Nemirovski · 2000
Earlier work this paper cites.
Lectures on modern convex optimization: analysis, algorithms, and engineering applications , volume 2
Aharon Ben-Tal and Arkadi Nemirovski · 2001
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
D. de Farias and B. Van Roy · 2003
Earlier work this paper cites.
Prox-method with rate of convergence O(1/t) for variational inequalities with lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski · 2004
Earlier work this paper cites.
Robust dynamic programming
G. Iyengar · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition probabilities
A. Nilim and L. El Ghaoui · 2005
Earlier work this paper cites.
Dynamic Programming and Optimal Control , volume 2
Dimitri Bertsekas · 2007
Earlier work this paper cites.
Naturalgradient actor-critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2007
Earlier work this paper cites.
Efficient projections onto the L-1 ball for learning in high dimensions
John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra · 2008
Earlier work this paper cites.
Percentile optimization for markov decision processes with parameter uncertainty
E. Delage and S. Mannor · 2010
Cited alongside, same era.
Optimization-based approximate dynamic programming
Marek Petrik · 2010
Cited alongside, same era.
A first-order primal-dual algorithm for convex problems with applications to imaging
Antonin Chambolle and Thomas Pock · 2011
Cited alongside, same era.
Proximal splitting methods in signal processing
Patrick L Combettes and Jean-Christophe Pesquet · 2011
Cited alongside, same era.
First order methods for nonsmooth convex large-scale optimization
Anatoli Juditsky, Arkadi Nemirovski, et al · 2011
Cited alongside, same era.
Approximate dynamic programming finally performs well in the game of Tetris
Victor Gabillon, Mohammad Ghavamzadeh, and Bruno Scherrer · 2013
Cited alongside, same era.
Anderson acceleration for reinforcement learning
Matthieu Geist and Bruno Scherrer · 2018
Later among the works it cites.
Data uncertainty in Markov chains: Application to cost-effectiveness analyses of medical innovations
Joel Goh, Mohsen Bayati, Stefanos A Zenios, Sundeep Singh, and David Moore · 2018
Later among the works it cites.
Robust Markov decision process: Beyond rectangularity
Vineet Goyal and Julien Grand-Clement · 2018
Later among the works it cites.
Fast Bellman updates for Robust MDPs
C.P. Ho, M. Petrik, and W.Wiesemann · 2018
Later among the works it cites.
Faster algorithms for extensive-form game solving via improved smoothing functions
Christian Kroer, Kevin Waugh, Fatma Kılınç-Karzan, and Tuomas Sandholm · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kullback-leibler divergence constrained distributionally robust optimization
Zhaolin Hu and L Jeff Hong · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
Robust Markov decision processes
W. Wiesemann, D. Kuhn, and B. Rustem · 2013
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of Tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Cited alongside, same era.
On the ergodic convergence rates of a first-order primal–dual algorithm
Antonin Chambolle and Thomas Pock · 2016
Cited alongside, same era.
Difference of convex functions programming applied to control with expert data
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2016
Cited alongside, same era.
Multi-model Markov decision processes
Lauren N Steimle, David L Kaufman, and Brian T Denton · 2018
Later among the works it cites.
Globally convergent type-I Anderson acceleration for non-smooth fixed-point iterations
Junzi Zhang, Brendan O’Donoghue, and Stephen Boyd · 2018
Later among the works it cites.
Probabilistic guarantees in robust optimization
Dimitris Bertsimas, Dick den Hertog, and Jean Pauphilet · 2019
Later among the works it cites.
Increasing iterate averaging for solving saddle-point problems
Yuan Gao, Christian Kroer, and Donald Goldfarb · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
A first-order approach to accelerated value iteration
Vineet Goyal and Julien Grand-Clement · 2019
Later among the works it cites.
Exploration bonus for regret minimization in discrete and continuous average reward mdps
QIAN Jian, Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2019
Later among the works it cites.
Active exploration in markov decision processes
Jean Tarbouriech and Alessandro Lazaric · 2019
Later among the works it cites.
Robust policies for proactive ICU transfers
Julien Grand-Clement, Carri W Chan, Vineet Goyal, and Gabriel Escobar · 2020
Closest in time.