Fetching the paper…
Reading the bibliography…
We propose a general framework for entropy-regularized average-reward reinforcement learning in Markov decision processes (MDPs).
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Dynamic Programming and Markov Processes
R. A. Howard · 1960
Earlier work this paper cites.
Risk-sensitive markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
Monotone Operators and the Proximal Point Algorithm
R. Tyrrell Rockafellar · 1976
Earlier work this paper cites.
Perturbation des méthodes d’optimisation. applications
B. Martinet · 1978
Earlier work this paper cites.
Modified policy iteration algorithms for discounted Markov decision processes
Martin L. Puterman and Moon Chirl Shin · 1978
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. Nemirovski and D. Yudin · 1983
Earlier work this paper cites.
Aggregating strategies
V. Vovk · 1990
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
The weighted majority algorithm
N. Littlestone and M. Warmuth · 1994
Earlier work this paper cites.
Markov Decision Processes
Martin L. Puterman · 1994
Earlier work this paper cites.
A Generalized Reinforcement Learning Model: Convergence and applications
M.L. Littman and Cs. Szepesvári · 1996
Earlier work this paper cites.
Optimal adaptive policies for Markov Decision Processes
A. N. Burnetas and M. N. Katehakis · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Risk sensitive markov decision processes
Steven I Marcus, Emmanual Fernández-Gaucherand, Daniel Hernández-Hernandez, Stefano Coraluppi, and Pedram Fard · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Relative loss bounds for multidimensional regression problems
J. Kivinen and M. Warmuth · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, and Fernando Pereira · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Cited alongside, same era.
A convergent form of approximate policy iteration
Theodore J. Perkins and Doina Precup · 2002
Cited alongside, same era.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Cited alongside, same era.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
Dynamic Programming and Optimal Control , volume 2
D. P. Bertsekas · 2007
Cited alongside, same era.
Stochastic Learning and Optimization: A Sensitivity-Based Approach
Xi-Ren Cao · 2007
Cited alongside, same era.
Dynamic policy programming
Mohammad Gheshlaghi Azar, Vicenç Gómez, and Hilbert J Kappen · 2012
Later among the works it cites.
An approximate solution method for large risk-averse markov decision processes
Marek Petrik and Dharmashankar Subramanian · 2012
Later among the works it cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Later among the works it cites.
Approximate modified policy iteration
B. Scherrer, V. Gabillon, M. Ghavamzadeh, and M. Geist · 2012
Later among the works it cites.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2012
Later among the works it cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Even-Dar, S. M. Kakade, and Y. Mansour · 2009
Cited alongside, same era.
Composite objective mirror descent
John C Duchi, Shai Shalev-Shwartz, Yoram Singer, and Ambuj Tewari · 2010
Cited alongside, same era.
The online loop-free stochastic shortest-path problem
G. Neu, A. György, and Cs. Szepesvári · 2010
Cited alongside, same era.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Cited alongside, same era.
Risk-averse dynamic programming for Markov decision processes
Andrzej Ruszczyński · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Cs. Szepesvári · 2010
Cited alongside, same era.
A. Zimin and G. Neu · 2013
Later among the works it cites.
Online learning in markov decision processes with changing cost sequences
T. Dick, A. György, and Cs · 2014
Later among the works it cites.
A survey of algorithms and analysis for adaptive online learning
H Brendan McMahan · 2014
Later among the works it cites.
Online Markov decision processes under bandit feedback
Gergely Neu, András György, Csaba Szepesvári, and András Antos · 2014
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
A new softmax operator for reinforcement learning
Kavosh Asadi and Michael L. Littman · 2016
Later among the works it cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Later among the works it cites.
Introduction to online convex optimization
Elad Hazan et al · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Guided policy search via approximate mirror descent
William H Montgomery and Sergey Levine · 2016
Later among the works it cites.
Fast rates for online learning in Linearly Solvable Markov Decision Processes
Gergely Neu and Vicenç Gómez · 2017
Closest in time.
PGQ: Combining policy gradient and Q-learning
Brendan O’Donoghue, Remi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2017
Closest in time.