Fetching the paper…
Reading the bibliography…
Modern deep neural networks are typically highly overparameterized.
Sharp Minima Can Generalize For Deep Nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 1938
Earlier work this paper cites.
Model Compression
Bucilua, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Predicting Parameters in Deep Learning
Denil, M., Shakibi, B., Dinh, L., Ranzato, M., and de Freitas, N · 2013
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2013
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2014
Earlier work this paper cites.
Speeding up Convolutional Neural Networks with Low Rank Expansions
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Rigid-motion scattering for image classification
Sifre, L. and Mallat, S · 2014
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Yang, Z., Moczulski, M., Denil, M., de Freitas, N., Smola, A., Song, L., and Wang, Z · 2014
Earlier work this paper cites.
Unitary Evolution Recurrent Neural Networks
Arjovsky, M., Shah, A., and Bengio, Y · 2015
Earlier work this paper cites.
Compressing Neural Networks with the Hashing Trick
Chen, W., Wilson, J. T., Tyree, S., Weinberger, K. Q., and Chen, Y · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Fast ConvNets Using Group-wise Brain Damage
Lebedev, V. and Lempitsky, V · 2015
Earlier work this paper cites.
Lin, M., Chen, Q., and Yan, S · 2015
Earlier work this paper cites.
ACDC: A Structured Efficient Linear Layer
Moczulski, M., Denil, M., Appleyard, J., and de Freitas, N · 2015
Earlier work this paper cites.
Structured Transforms for Small-Footprint Deep Learning
Sindhwani, V., Sainath, T. N., and Kumar, S · 2015
Cited alongside, same era.
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Cited alongside, same era.
An empirical analysis of the optimization of deep network loss surfaces
Im, D. J., Tao, M., and Branson, K · 2016
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Exploring Sparsity in Recurrent Neural Networks
Narang, S., Elsen, E., Diamos, G., and Sengupta, S · 2017
Later among the works it cites.
Theory of Deep Learning III: explaining the non-overfitting puzzle
Poggio, T., Kawaguchi, K., Liao, Q., Miranda, B., Rosasco, L., Boix, X., Hidary, J., and Mhaskar, H · 2017
Later among the works it cites.
Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
Wu, L., Zhu, Z., and E, W · 2017
Later among the works it cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Later among the works it cites.
Nullhop: A flexible convolutional neural network accelerator based on sparse representations of feature maps
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Cited alongside, same era.
Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Learning structured sparsity in deep neural networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Deep Rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R · 2017
Cited alongside, same era.
SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
NeST: A Neural Network Synthesis Tool Based on a Grow-and-Prune Paradigm
Dai, X., Yin, H., and Jha, N. K · 2017
Cited alongside, same era.
Aimar, A., Mostafa, H., Calabrese, E., Rios-Navarro, A., Tapiador-Morales, R., Lungu, I.-A., Milde, M. B., Corradi, F., Linares-Barranco, A., Liu, S.-C., et al · 2018
Later among the works it cites.
The loss landscape of overparameterized neural networks
Cooper, Y · 2018
Later among the works it cites.
Grow and Prune Compact, Fast, and Accurate LSTMs
Dai, X., Yin, H., and Jha, N. K · 2018
Later among the works it cites.
The Lottery Ticket Hypothesis: Finding Small, Trainable Neural Networks
Frankle, J. and Carbin, M · 2018
Later among the works it cites.
Measuring the Intrinsic Dimension of Objective Landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Later among the works it cites.
Rethinking the Value of Network Pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Later among the works it cites.
Training wide residual networks for deployment using a single bit for each weight
McDonnell, M. D · 2018
Later among the works it cites.
Recovering from Random Pruning: On the Plasticity of Deep Convolutional Neural Networks
Mittal, D., Bhardwaj, S., Khapra, M. M., and Ravindran, B · 2018
Later among the works it cites.
Sensitivity and Generalization in Neural Networks: an Empirical Study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Network Compression using Correlation Analysis of Layer Responses
Suau, X., Zappella, L., and Apostoloff, N · 2018
Later among the works it cites.
Learning invariance with compact transforms
Thomas, A. T., Gu, A., Dao, T., Rudra, A., and Christopher, R · 2018
Later among the works it cites.
Low-Cost Parameterizations of Deep Convolution Neural Networks
Treister, E., Ruthotto, L., Sharoni, M., Zafrani, S., and Haber, E · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2041
Closest in time.