Fetching the paper…
Reading the bibliography…
This work presents a new algorithm for training recurrent neural networks (although ideas are applicable to feedforward networks as well).
The Heat Equation
Widder, D. V. (1975) · 1975
Earlier work this paper cites.
A Connectionist Machine for Genetic Hillclimbing
Ackley, D. (1987) · 1987
Earlier work this paper cites.
A method to convexify functions via curve evolution
Vese, L. (1999) · 1999
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2007
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
Martens, J. and Sutskever, I. (2011) · 2011
Earlier work this paper cites.
Randomized smoothing for stochastic optimization
Duchi, J. C., Bartlett, P. L., and Wainwright, M. J. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Optimization by Gaussian Smoothing with Application to Geometric Alignment
Mobahi, H. (2012) · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Cited alongside, same era.
Learning with pseudo-ensembles
Bachman, P., Alsharif, O., and Precup, D. (2014) · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Marginalized denoising auto-encoders for nonlinear representations
Chen, M., Weinberger, K. Q., Sha, F., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Generalization bounds for neural networks through tensor factorization
Janzamin, M., Sedghi, H., and Anandkumar, A. (2015) · 2015
Later among the works it cites.
On the Link Between Gaussian Homotopy Continuation and Convex Envelopes
Mobahi, H. and Fisher III, J. W. (2015) · 2015
Later among the works it cites.
Adding gradient noise improves learning for very deep networks
Neelakantan, A., Vilnis, L., Le, Q. V., Sutskever, I., Kaiser, L., Kurach, K., and Martens, J. (2015) · 2015
Later among the works it cites.
Annealed gradient descent for deep learning
Pan, H. and Jiang, H. (2015) · 2015
Later among the works it cites.
On the quality of the initial basin in overspecified neural networks
Safran, I. and Shamir, O. (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pascanu, R., Dauphin, Y. N., Ganguli, S., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2015) · 2015
Cited alongside, same era.
On graduated optimization for stochastic non-convex problems
Hazan, E., Levy, K. Y., and Shalev-Shwartz, S. (2015) · 2015
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G. E., Deng, L., Yu, D., Dahl, G. E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., and Kingsbury, B. (2012a)
Cited in the paper.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012b)
Cited in the paper.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. (2015) · 2015
Later among the works it cites.
ℓ 1 \ell_{1} -regularized neural networks are improperly learnable in polynomial time
Zhang, Y., Lee, J. D., and Jordan, M. I. (2015) · 2015
Later among the works it cites.
Closed form for some gaussian convolutions
Mobahi, H. (2016) · 2016
Closest in time.