Fetching the paper…
Reading the bibliography…
Deep artificial neural networks (DNNs) are typically trained via gradient-based learning algorithms, namely backpropagation.
A stochastic approximation method
Robbins H. and Monro S · 1951
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Hoerl A. E. and Kennard R. W · 1970
Earlier work this paper cites.
Genetic algorithms
Holland J. H · 1992
Earlier work this paper cites.
Q-learning
Watkins C. J. and Dayan P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams R. J · 1992
Earlier work this paper cites.
On the effectiveness of crossover in simulated evolutionary optimization
Fogel D. B. and Stayton L. C · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter S. and Schmidhuber J · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton R. S. and Barto A. G · 1998
Earlier work this paper cites.
Introduction to evolutionary computing , volume 53
Eiben A. E., Smith J. E., and others · 2003
Earlier work this paper cites.
Practical genetic algorithms
Haupt R. L. and Haupt S. E · 2004
Earlier work this paper cites.
Compositional pattern producing networks: A novel abstraction of development
Stanley K. O · 2007
Earlier work this paper cites.
Natural evolution strategies
Wierstra D., Schaul T., Peters J., and Schmidhuber J · 2008
Earlier work this paper cites.
Overcoming the bootstrap problem in evolutionary robotics using behavioral diversity
Mouret J.-B. and Doncieux S · 2009
Earlier work this paper cites.
A hypercube-based indirect encoding for evolving large-scale neural networks
Stanley K. O., D’Ambrosio D. B., and Gauci J · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot X. and Bengio Y · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke F., Osendorfer C., Rückstieß T., Graves A., Peters J., and Schmidhuber J · 2010
Earlier work this paper cites.
On the performance of indirect encoding across the continuum of regularity
Clune J., Stanley K. O., Pennock R. T., and Ofria C · 2011
Earlier work this paper cites.
Conversational speech transcription using context-dependent deep neural networks
Seide F., Li G., and Yu D · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky A., Sutskever I., and Hinton G. E · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov E., Erez T., and Tassa Y · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare M. G., Naddaf Y., Veness J., and Bowling M · 2013
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Cho K., Van Merriënboer B., Bahdanau D., and Bengio Y · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma D. and Ba J · 2014
Cited alongside, same era.
On the saddle point problem for non-convex optimization
Pascanu R., Dauphin Y. N., Ganguli S., and Bengio Y · 2014
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih V., Badia A. P., Mirza M., Graves A., Lillicrap T., Harley T., Silver D., and Kavukcuoglu K · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband I., Blundell C., Pritzel A., and Van Roy B · 2016
Later among the works it cites.
Pugh J. K., Soros L. B., and Stanley K. O
2016
Later among the works it cites.
Improved techniques for training gans
Salimans T., Goodfellow I., Zaremba W., Cheung V., Radford A., and Chen X · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt H., Guez A., and Silver D · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare M. G., Dabney W., and Munos R · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout: A simple way to prevent neural networks from overfitting
Srivastava N., Hinton G., Krizhevsky A., Sutskever I., and Salakhutdinov R · 2014
Cited alongside, same era.
Robots that can adapt like animals
Cully A., Clune J., Tarapore D., and Mouret J.-B · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He K., Zhang X., Ren S., and Sun J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe S. and Szegedy C · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih V., Kavukcuoglu K., Silver D., Rusu A. A., Veness J., Bellemare M. G., Graves A., Riedmiller M., Fidjeland A. K., Ostrovski G., and others · 2015
Cited alongside, same era.
Illuminating search spaces by mapping elites
Mouret J. and Clune J · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Nair A., Srinivasan P., Blackwell S., Alcicek C., Fearon R., De Maria A., Panneershelvam V., Suleyman M., Beattie C., Petersen S., and others · 2015
Cited alongside, same era.
Conti E., Madhavan V., Petroski Such F., Lehman J., Stanley K. O., and Clune J · 2017
Closest in time.
Noisy networks for exploration
Fortunato M., Azar M. G., Piot B., Menick J., Osband I., Graves A., Mnih V., Munos R., Hassabis D., Pietquin O., and others · 2017
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
Hessel M., Modayil J., Van Hasselt H., Schaul T., Ostrovski G., Dabney W., Horgan D., Piot B., Azar M., and Silver D · 2017
Closest in time.
Self-normalizing neural networks
Klambauer G., Unterthiner T., Mayr A., and Hochreiter S · 2017
Closest in time.
ES is more than just a traditional finite-difference approximator
Lehman J., Chen J., Clune J., and Stanley K. O · 2017
Closest in time.
Hierarchical representations for efficient architecture search
Liu H., Simonyan K., Vinyals O., Fernando C., and Kavukcuoglu K · 2017
Closest in time.
Miikkulainen R., Liang J., Meyerson E., Rawal A., Fink D., Francon O., Raju B., Navruzyan A., Duffy N., and Hodjat B · 2017
Closest in time.
Parameter space noise for exploration
Plappert M., Houthooft R., Dhariwal P., Sidor S., Chen R. Y., Chen X., Asfour T., Abbeel P., and Andrychowicz M · 2017
Closest in time.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Salimans T., Ho J., Chen X., Sidor S., and Sutskever I · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans T., Ho J., Chen X., and Sutskever I · 2017
Closest in time.
Proximal policy optimization algorithms
Schulman J., Wolski F., Dhariwal P., Radford A., and Klimov O · 2017
Closest in time.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu Y., Mansimov E., Grosse R. B., Liao S., and Ba J · 2017
Closest in time.