Fetching the paper…
Reading the bibliography…
Background: Deep learning models are typically trained using stochastic gradient descent or one of its variants.
Random Walk in a Random Environment and 1f Noise
Marinari, E., Parisi, G., Ruelle, D., and Windey, P · 1983
Earlier work this paper cites.
Multidimensional random walks in random environments with subclassical limiting behavior
Durrett, R · 1986
Earlier work this paper cites.
Anomalous diffusion in random media of any dimensionality
Bouchaud, J. P. and Comtet, A · 1987
Earlier work this paper cites.
Anomalous diffusion in disordered media: statistical mechanisms, models and physical applications
Bouchaud, J. P. and Georges, A · 1990
Earlier work this paper cites.
Regularization theory and neural networks architectures
Girosi, F., Jones, M., and Poggio, T · 1995
Earlier work this paper cites.
Statistics of critical points of Gaussian fields on large-dimensional spaces
Bray, A. J. and Dean, D. S · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., et al · 2012
Earlier work this paper cites.
Neural Networks: Tricks of the Trade
Montavon, G., Orr, G., and Müller, K.-R · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., LeCun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., and Gulcehre, C · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
Li, M., Zhang, T., Chen, Y., and Smola, A. J · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. e. a · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Later among the works it cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Later among the works it cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., et al · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Amodei, D., Anubhai, R., Battenberg, E., et al · 2015
Cited alongside, same era.
The Loss Surfaces of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Cited alongside, same era.
Dauphin, Y., de Vries, H., Chung, J., and Bengio, Y · 2015
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Cited alongside, same era.
Wu, Y., Schuster, M., Chen, Z., et al · 2016
Later among the works it cites.
Wide residual networks
Zagoruyko, K · 2016
Later among the works it cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., et al · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Closest in time.
Scaling Distributed Machine Learning with System and Algorithm Co-design
Li, M · 2017
Closest in time.
The Implicit Bias of Gradient Descent on Separable Data
Soudry, D., Hoffer, E., and Srebro, N · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E · 2017
Closest in time.
Scaling sgd batch size to 32k for imagenet training
You, Y., Gitman, I., and Ginsburg, B · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Closest in time.