Understand
In this paper we revisit the method of off-policy corrections for reinforcement learning (COP-TD) pioneered by Hallak et al.
- (2017).
- Under this method, online updates to the value function are reweighted to avoid divergence issues typical of off-policy learning.
- While Hallak et al.'s solution is appealing, it cannot easily be transferred to nonlinear function approximation.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…