Fetching the paper…
Reading the bibliography…
Most learning algorithms are not invariant to the scale of the function that is being approximated.
A logical calculus of the ideas immanent in nervous activity
W. S. McCulloch and W. Pitts · 1943
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Principles of Neurodynamics
F. Rosenblatt · 1962
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Asymmetric least squares estimation and testing
W. K. Newey and J. L. Powell · 1987
Earlier work this paper cites.
Regression percentiles using asymmetric squared error loss
B. Efron · 1991
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. I. Amari · 1998
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
S. Hochreiter · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
H. J. Kushner and G. Yin · 2003
Earlier work this paper cites.
Double Q-learning
H. van Hasselt · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
J. S. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Insights in Reinforcement Learning
H. van Hasselt · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. B. Grosse · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
J. Schmidhuber · 2015
Later among the works it cites.
Learning from delayed rewards
C. J. C. H. Watkins · 2015
Later among the works it cites.
Increasing the action gap: New operators for reinforcement learning
M. G. Bellemare, G. Ostrovski, A. Guez, P. S. Thomas, and R. Munos · 2016
Closest in time.
State of the art control of atari games using shallow reinforcement learning
Y. Liang, M. C. Machado, E. Talvitie, and M. H. Bowling · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Normalized online learning
S. Ross, P. Mineiro, and J. Langford · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, and K. Kavukcuoglu · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
A method for stochastic optimization
D. P. Kingma and J. B. Adam · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Closest in time.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Closest in time.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Closest in time.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Closest in time.
Deep reinforcement learning with Double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Closest in time.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, T. Schaul, M. Hessel, H. van Hasselt, and M. Lanctot · 2016
Closest in time.