A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Original
Cao, Y · 1902
Earlier work this paper cites.
Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation
Original
Yang, G · 1902
Earlier work this paper cites.
A selective overview of deep learning
Original
Fan, J · 1904
Earlier work this paper cites.
Striving for simplicity in off-policy deep reinforcement learning
Original
Agarwal, R · 1907
Earlier work this paper cites.
A fine-grained spectral perspective on neural networks
Original
Yang, G · 1907
Earlier work this paper cites.
Reinforcement learning in healthcare: A survey
Original
Yu, C · 1908
Earlier work this paper cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Original
Bai, Y · 1910
Earlier work this paper cites.
Benchmarking batch deep reinforcement learning algorithms
Original
Fujimoto, S · 1910
Earlier work this paper cites.
A finite-time analysis of Q-learning with neural network function approximation
Original
Xu, P · 1912
Earlier work this paper cites.
Theory of games and economic behavior
Von Neumann, J · 1947
Earlier work this paper cites.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N · 1958
Earlier work this paper cites.
Projection pursuit regression
Friedman, J. H · 1981
Earlier work this paper cites.
Optimal global rates of convergence for nonparametric regression
Stone, C. J · 1982
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Neural nets with superlinear VC-dimension
Maass, W · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J · 1996
Earlier work this paper cites.
Stochastic and shortest path games: Theory and algorithms
Patek, S. D · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: The size of the weights is more important than the size of the network
Bartlett, P. L · 1998
Earlier work this paper cites.
Almost linear VC dimension bounds for piecewise polynomial networks
Bartlett, P. L · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S · 2000
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Bowling, M · 2001
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Boyan, J. A · 2002
Earlier work this paper cites.
Why do deep residual networks generalize better than deep feedforward networks?–A neural tangent kernel perspective
Original
Huang, K · 2002
Earlier work this paper cites.
Value function approximation in zero-sum Markov games
Lagoudakis, M. G · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
Murphy, S. A · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D · 2005
Earlier work this paper cites.
A generalization error for Q-learning
Murphy, S. A · 2005
Earlier work this paper cites.
Neural fitted Q iteration – First experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Value-iteration based fitted policy iteration: Learning with a single trajectory
Antos, A · 2007
Earlier work this paper cites.
AWESOME: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
Conitzer, V · 2007
Earlier work this paper cites.
Kernel methods in machine learning
Hofmann, T · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A · 2008
Earlier work this paper cites.
Introduction to nonparametric estimation
Tsybakov, A. B · 2008
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Anthony, M · 2009
Earlier work this paper cites.
Regularized fitted Q-iteration for planning in continuous-space Markovian decision problems
Farahmand, A.-m · 2009
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A · 2009
Earlier work this paper cites.
Reinforcement learning design for cancer clinical trials
Zhao, Y · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A.-m · 2010
Earlier work this paper cites.