Fetching the paper…

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · Around