Fetching the paper…
Reading the bibliography…
We show that the simplest actor-critic method -- a linear softmax policy updated with TD through interaction with a linear MDP, but featuring no explicit regularization or exploration -- does not merely find an optimal policy, but moreover prefers high entropy optimal policies.
Ziwei Ji and Matus Telgarsky · 1909
Earlier work this paper cites.
On convergence proofs on perceptrons
Albert B.J. Novikoff · 1962
Earlier work this paper cites.
Conductance and the rapid mixing property for markov chains: The approximation of permanent resolved
Mark Jerrum and Alistair Sinclair · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
The mixing rate of markov chains, an isoperimetric inequality, and computing the volume
L. Lovasz and M. Simonovits · 1990
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Tong Zhang and Bin Yu · 2005
Earlier work this paper cites.
Markov chains and mixing times
David A. Levin, Yuval Peres, and Elizabeth L. Wilmer · 2006
Earlier work this paper cites.
Q-learning with linear function approximation
Francisco S Melo and M Isabel Ribeiro · 2007
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Cited alongside, same era.
Optimal Transport: Old and New
Cèdric Villani · 2008
Cited alongside, same era.
Adaptive ε \varepsilon -greedy exploration in reinforcement learning based on value differences
Michel Tokic · 2010
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2011
Cited alongside, same era.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Cited alongside, same era.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Cited alongside, same era.
Convex optimization: Algorithms and complexity
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and Lei Ying · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
Finite-sample analysis for SARSA with linear function approximation
Shaofeng Zou, Tengyu Xu, and Yingbin Liang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sébastien Bubeck · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Yanli Liu, Kaiqing Zhang, Tamer Basar, and Wotao Yin · 2020
Later among the works it cites.
Gradient methods never overfit on separable data
Ohad Shamir · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Tengyu Xu, Zhe Wang, and Yingbin Liang · 2020
Later among the works it cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M. Kakade, and Wen Sun · 2021
Closest in time.
Finite sample analysis of two-time-scale natural actor-critic algorithm
Sajad Khodadadian, Thinh T Doan, Siva Theja Maguluri, and Justin Romberg · 2021
Closest in time.
Introduction to reinforcement learning: Lecture 7
David Silver · 2021
Closest in time.