Fetching the paper…

Stop Regressing: Training Value Functions via Classification for Scalable Deep RL · Around