Fetching the paper…
Reading the bibliography…
Training neural networks involves finding minima of a high-dimensional non-convex loss function.
Minimum spanning trees and single linkage cluster analysis
Gower, J. C. and Ross, G. J. S · 1969
Earlier work this paper cites.
Neural network ensembles
Hansen, L. K. and Salamon, P · 1990
Earlier work this paper cites.
Nudged elastic band method for finding minimum energy paths of transitions
Jónsson, H., Mills, G., and Jacobsen, K. W · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Archetypal energy landscapes
Wales, D. J., Miller, M. A., and Walsh, T. R · 1998
Earlier work this paper cites.
Improved tangent estimate in the nudged elastic band method for finding minimum energy paths and saddle points
Henkelman, G. and Jónsson, H · 2000
Earlier work this paper cites.
Optimization methods for finding minimum energy paths
Sheppard, D., Terrell, R., and Henkelman, G · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Flexible, high performance convolutional neural networks for image classification
Ciresan, D. C., Meier, U., Masci, J., Maria Gambardella, L., and Schmidhuber, J · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A.-r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., et al · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r., and Hinton, G · 2013
Earlier work this paper cites.
The loss surface of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gülçehre, Ç., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
Montúfar, G., Pascanu, R., Cho, K., and Bengio, Y · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
An automated nudged elastic band method
Kolsbjerg, E. L., Groves, M. N., and Hammer, B · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y · 2016
Later among the works it cites.
Energy landscapes for machine learning
Ballard, A. J., Das, R., Martiniani, S., Mehta, D., Sagun, L., Stevenson, J. D., and Wales, D. J · 2017
Later among the works it cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Later among the works it cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Weinberger, K. Q., and van der Maaten, L · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vinyals, O. and Le, Q. V · 2015
Cited alongside, same era.
Energy landscapes for a machine learning application to series data
Ballard, A. J., Stevenson, J. D., Das, R., and Wales, D. J · 2016
Cited alongside, same era.
Topology and Geometry of Half-Rectified Network Optimization
Freeman, C. D. and Bruna, J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Li, H., Xu, Z., Taylor, G., and Goldstein, T · 2017
Later among the works it cites.
Learning efficient convolutional networks through network slimming
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Later among the works it cites.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Later among the works it cites.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
Sagun, L., Evci, U., Ugur Guney, V., Dauphin, Y., and Bottou, L · 2017
Later among the works it cites.
Toward human parity in conversational speech recognition
Xiong, W., Droppo, J., Huang, X., Seide, F., Seltzer, M. L., Stolcke, A., Yu, D., and Zweig, G · 2017
Later among the works it cites.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Closest in time.