Fetching the paper…
Reading the bibliography…
Markov decision processes (MDPs) are used to model stochastic systems in many applications.
A lipschitz condition preserving extension for a vector function
Frederick Albert Valentine · 1945
Earlier work this paper cites.
Two observations about the method of succesive approximations
MA Krasnoselskii · 1955
Earlier work this paper cites.
Dynamic programming and markov processes
Ronald A Howard · 1960
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Accelerated procedures for the solution of discrete Markov control problems
H Kushner and A Kleinman · 1971
Earlier work this paper cites.
Accelerated computation of the expected discounted return in a Markov chain
Evan L Porteus and John C Totten · 1978
Earlier work this paper cites.
On the convergence of policy iteration in stationary dynamic programming
Martin L Puterman and Shelby L Brumelle · 1979
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
On the algorithm of Pollatschek and Avi-ltzhak
Jerzy A Filar and Boleslaw Tolwinski · 1991
Earlier work this paper cites.
Random walks and the effective resistance of networks
Prasad Tetali · 1991
Earlier work this paper cites.
On the generation of Markov decision processes
TW Archibald, KIM McKinnon, and LC Thomas · 1995
Earlier work this paper cites.
A k-step look-ahead analysis of value iteration algorithms for Markov decision processes
Meir Herzberg and Uri Yechiali · 1996
Earlier work this paper cites.
Application of stochastic dynamic programming to optimal fire management of a spatially structured threatened species
Hugh Possingham and G Tuck · 1997
Earlier work this paper cites.
The Lyapunov exponent and joint spectral radius of pairs of matrices are hard—when not impossible—to compute and to approximate
John N Tsitsiklis and Vincent D Blondel · 1997
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
D. de Farias and B. Van Roy · 2003
Earlier work this paper cites.
Computationally efficient approximations of the joint spectral radius
Vincent D Blondel and Yurii Nesterov · 2005
Earlier work this paper cites.
Robust dynamic programming
G. Iyengar · 2005
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
The joint spectral radius: theory and applications , volume 385
Raphaël Jungers · 2009
Earlier work this paper cites.
Optimization-based approximate dynamic programming
Marek Petrik · 2010
Cited alongside, same era.
Acceleration operators in the value iteration algorithms for Markov decision processes
Oleksandr Shlakhter, Chi-Guhn Lee, Dmitry Khmelev, and Nasser Jaber · 2010
Cited alongside, same era.
The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate
Y. Ye · 2011
Cited alongside, same era.
Markov chains: Gibbs fields, Monte Carlo simulation, and queues , volume 31
Pierre Brémaud · 2013
Cited alongside, same era.
Reversible Markov decision processes with an average-reward criterion
Randy Cogill and Cheng Peng · 2013
Cited alongside, same era.
Smooth manifolds
John M Lee · 2013
Cited alongside, same era.
Markov decision processes in practice
Richard J Boucherie and Nico M Van Dijk · 2017
Later among the works it cites.
Anderson acceleration for reinforcement learning
Matthieu Geist and Bruno Scherrer · 2018
Later among the works it cites.
Data uncertainty in Markov chains: Application to cost-effectiveness analyses of medical innovations
Joel Goh, Mohsen Bayati, Stefanos A Zenios, Sundeep Singh, and David Moore · 2018
Later among the works it cites.
Robust Markov Decision Process: Beyond Rectangularity
V. Goyal and J. Grand-Clément · 2018
Later among the works it cites.
Fast Bellman updates for robust MDPs
C.P. Ho, M. Petrik, and W.Wiesemann · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust Markov decision processes
W. Wiesemann, D. Kuhn, and B. Rustem · 2013
Cited alongside, same era.
Fast alternating direction optimization methods
Tom Goldstein, Brendan O’Donoghue, Simon Setzer, and Richard Baraniuk · 2014
Cited alongside, same era.
Markov Decision Processes : Discrete Stochastic Dynamic Programming
M.L. Puterman · 2014
Cited alongside, same era.
Markov Decision Process (MDP) toolbox for python
Steven Cordwell, Yasser Gonzalez, and Theja Tulabandhula · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
Adaptive restart for accelerated gradient schemes
Brendan O’donoghue and Emmanuel Candes · 2015
Cited alongside, same era.
Vladimir Yu Protasov and Aleksandar Cvetković · 2018
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving Markov decision processes
Aaron Sidford, Mengdi Wang, Xian Wu, and Yinyu Ye · 2018
Later among the works it cites.
Temporal regularization for Markov decision process
Pierre Thodoroff, Audrey Durand, Joelle Pineau, and Doina Precup · 2018
Later among the works it cites.
Globally convergent type-I Anderson acceleration for non-smooth fixed-point iterations
Junzi Zhang, Brendan O’Donoghue, and Stephen Boyd · 2018
Later among the works it cites.
The operator approach to entropy games
Marianne Akian, Stéphane Gaubert, Julien Grand-Clément, and Jérémie Guillaud · 2019
Closest in time.
A generic online acceleration scheme for optimization algorithms via relaxation and inertia
Franck Iutzeler and Julien M Hendrickx · 2019
Closest in time.
Exploration bonus for regret minimization in discrete and continuous average reward MDPs
QIAN Jian, Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2019
Closest in time.
Active exploration in Markov decision processes
Jean Tarbouriech and Alessandro Lazaric · 2019
Closest in time.
On connections between constrained optimization and reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2019
Closest in time.
Stochastic approximation with cone-contractive operators: Sharp L-infty-bounds for Q-learning
Martin J Wainwright · 2019
Closest in time.
Robust policies for proactive ICU transfers
Julien Grand-Clément, Carri W Chan, Vineet Goyal, and Gabriel Escobar · 2020
Closest in time.
Julien Grand-Clément · 2021
Closest in time.
Taylor expansion of discount factors
Yunhao Tang, Mark Rowland, Rémi Munos, and Michal Valko · 2021
Closest in time.