Fetching the paper…

Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning · Around