SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Later among the works it cites.
Deep reinforcement learning: An overview
Original
Li, Y · 2017
Later among the works it cites.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
Schulman, J · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D · 2017
Later among the works it cites.
Boosted fitted Q-iteration
Tosatto, S · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Original
Bhandari, J · 2018
Later among the works it cites.
A note on lazy training in supervised differentiable programming
Original
Chizat, L · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Original
Du, S. S · 2018
Later among the works it cites.
Global convergence of policy gradient methods for linearized control problems
Original
Fazel, M · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Original
Haarnoja, T · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q · 2018
Later among the works it cites.
Convergent actor-critic algorithms under off-policy training and function approximation
Original
Maei, H. R · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Original
Malik, D · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Original
Tu, S · 2018
Later among the works it cites.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Original
Zou, D · 2018
Later among the works it cites.
Solving the Rubik’s cube with deep reinforcement learning and search
Agostinelli, F · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O · 2019
Later among the works it cites.
Alphastar: Mastering the Real-Time Strategy Game StarCraft II
Vinyals, O · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T · 2019
Later among the works it cites.
Two time-scale off-policy TD learning: Non-asymptotic analysis over Markovian samples
Xu, T · 2019
Later among the works it cites.
Finite-sample analysis for SARSA with linear function approximation
Zou, S · 2019
Later among the works it cites.