Fetching the paper…
Reading the bibliography…
We investigate the difficulties of training sparse neural networks and make new observations about optimization dynamics and the energy landscape within the sparse regime.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Statistics of critical points of Gaussian fields on large-dimensional spaces
Bray, A. J. and Dean, D. S · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
The Loss Surface of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2014
Earlier work this paper cites.
Identifying and Attacking the Saddle Point Problem in High-dimensional Non-convex Optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
https://ui.adsabs.harvard.edu/abs/2014arXiv1412.6615S
Sagun, L., Ugur Guney, V., Ben Arous, G., and LeCun, Y · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Bayesian compression for deep learning
Louizos, C., Ullrich, K., and Welling, M · 2017
Cited alongside, same era.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Later among the works it cites.
Efficient Neural Audio Synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A., Dieleman, S., and Kavukcuoglu, K · 2018
Later among the works it cites.
Rethinking the Value of Network Pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Later among the works it cites.
Deep residual learning for image steganalysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Cited alongside, same era.
http://arxiv.org/abs/1704.05119
Narang, S., Diamos, G. F., Sengupta, S., and Elsen, E · 2017
Cited alongside, same era.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
Sagun, L., Evci, U., Güney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
Deep Rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R. A · 2018
Cited alongside, same era.
Learning Sparse Neural Networks through L 0 L_{0} Regularization
Christos Louizos, Max Welling, D. P. K · 2018
Cited alongside, same era.
The Lottery Ticket Hypothesis: Training Pruned Neural Networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Wu, S., Zhong, S., and Liu, Y · 2018
Later among the works it cites.
To Prune, or Not to Prune: Exploring the Efficacy of Pruning for Model Compression
Zhu, M. and Gupta, S · 2018
Later among the works it cites.
The Lottery Ticket Hypothesis at Scale
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2019
Closest in time.
The State of Sparsity in Deep Neural Networks
Gale, T., Elsen, E., and Hooker, S · 2019
Closest in time.
Mostafa, H. and Wang, X · 2019
Closest in time.
Evolving and Understanding Sparse Deep Neural Networks using Cosine Similarity
Pieterse, J. and Mocanu, D. C · 2019
Closest in time.