Fetching the paper…
Reading the bibliography…
This is a short communication on a Lyapunov function argument for softmax in bandit problems.
Optimality and approximation with policy gradient methods in markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2019
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
J. Bhandari and D. Russo · 2019
Earlier work this paper cites.
Regret analysis of a markov policy gradient algorithm for multi-arm bandits
D. Denisov and N. Walton · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…