Analysis of classification-based policy iteration algorithms
Lazaric, A · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V · 2016
Later among the works it cites.
Expressiveness of rectifier networks
Pan, X · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D · 2016
Later among the works it cites.
SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T · 2017
Later among the works it cites.
Deep reinforcement learning: An overview
Original
Li, Y · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
Schulman, J · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y · 2017
Later among the works it cites.
A note on lazy training in supervised differentiable programming
Original
Chizat, L · 2018
Later among the works it cites.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Original
Espeholt, L · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Original
Fazel, M · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Deep neural networks as Gaussian processes
Lee, J · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Original
Malik, D · 2018
Later among the works it cites.
Lectures on Convex Optimization
Nesterov, Y · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Original
Papini, M · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S · 2018
Later among the works it cites.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Original
Tu, S · 2018
Later among the works it cites.
Deep reinforcement learning for NLP
Wang, W. Y · 2018
Later among the works it cites.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Original
Zou, D · 2018
Later among the works it cites.
Hessian aided policy gradient
Shen, Z · 2019
Closest in time.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
Vinyals, O · 2019
Closest in time.