Fetching the paper…
Reading the bibliography…
In this short note we derive a relationship between the Bregman divergence from the current policy to the optimal policy and the suboptimality of the current value function in a regularized Markov decision process.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Primer on monotone operator methods
E. K. Ryu and S. Boyd · 2016
Earlier work this paper cites.
On the global convergence rates of softmax policy gradient methods
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Cited alongside, same era.
Making sense of reinforcement learning and probabilistic inference
B. O’Donoghue, I. Osband, and C. Ionescu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…