Fetching the paper…
Reading the bibliography…
Dynamic Programming (DP) provides standard algorithms to solve Markov Decision Processes.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Yurii Nesterov · 2009
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Local policy search in a convex space and conservative policy iteration as boosted policy search
Bruno Scherrer and Matthieu Geist · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Softened approximate policy iteration for markov games
Julien Pérolat, Bilal Piot, Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2016
Cited alongside, same era.
Is the bellman residual a bad proxy?
Matthieu Geist, Bilal Piot, and Olivier Pietquin · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2019
Closest in time.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Closest in time.
A first-order approach to accelerated value iteration
Vineet Goyal and Julien Grand-Clement · 2019
Closest in time.
Politex: Regret bounds for policy iteration using expert prediction
Nevena Lazic, Yasin Abbasi-Yadkori, Kush Bhatia, Gellert Weisz, Peter Bartlett, and Csaba Szepesvari · 2019
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep conservative policy iteration
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2019
Closest in time.