2017

f-Divergence constrained policy improvement

Belousov, Boris, Peters, Jan

Understand

To ensure stability of learning, state-of-the-art generalized policy iteration algorithms augment the policy improvement step with a trust region constraint bounding the information loss.

  • The size of the trust region is commonly determined by the Kullback-Leibler (KL) divergence, which not only captures the notion of distance well but also yields closed-form solutions.
  • In this paper, we consider a more general class of f-divergences and derive the corresponding policy update rules.
  • The generic solution is expressed through the derivative of the convex conjugate function to f and includes the KL solution as a special case.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…