Fetching the paper…

Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima · Around