2019

Off-Policy Deep Reinforcement Learning by Bootstrapping the Covariate Shift

Gelada, Carles, Bellemare, Marc G.

Understand

In this paper we revisit the method of off-policy corrections for reinforcement learning (COP-TD) pioneered by Hallak et al.

  • (2017).
  • Under this method, online updates to the value function are reweighted to avoid divergence issues typical of off-policy learning.
  • While Hallak et al.'s solution is appealing, it cannot easily be transferred to nonlinear function approximation.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…