Fetching the paper…
Reading the bibliography…
Overparameterized Neural Networks (NN) display state-of-the-art performance.
Yang, G. (2019) · 1902
Earlier work this paper cites.
Mean-field behaviour of neural tangent kernel for deep neural networks
Hayou, S., A. Doucet, and J. Rousseau (2020) · 1905
Earlier work this paper cites.
Large scale learning of general visual representations for transfer
Kolesnikov, A., L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby (2019) · 1912
Earlier work this paper cites.
La distribution de la plus grande de n n valeurs
Von Mises, R. (1936) · 1936
Earlier work this paper cites.
Inequalities
Hardy, G., J. Littlewood, and G. Pólya (1952) · 1952
Earlier work this paper cites.
Limit theorems for random central order statistics
Puri, M. and S. Ralescu (1986) · 1986
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Mozer, M. and P. Smolensky (1989) · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., J. Denker, and S. Solla (1990) · 1990
Earlier work this paper cites.
Convex Functions, Partial Orderings, and Statistical Applications
Pečarić, J., F. Proschan, and Y. Tong (1992) · 1992
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B., D. Stork, and W. Gregory (1993) · 1993
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. (1995) · 1995
Earlier work this paper cites.
Dhp: Differentiable meta pruning via hypernetworks
Li, Y., S. Gu, K. Zhang, L. Van Gool, and R. Timofte (2020) · 2003
Earlier work this paper cites.
Pruning neural networks at initialization: Why are we missing the mark?
Frankle, J., G. Dziugaite, D. Roy, and M. Carbin (2020) · 2009
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., X. Zhang, S. Ren, and J. Sun (2015) · 2015
Earlier work this paper cites.
Random synaptic feedback weights support error backpropagation for deep learning
Lillicrap, T., D. Cownden, D. Tweed, and C. Akerman (2016) · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli (2016) · 2016
Cited alongside, same era.
Probability in High Dimension
Van Handel, R. (2016) · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., S. Bengio, M. Hardt, B. Recht, and O. Vinyals (2016) · 2016
Cited alongside, same era.
Compression-aware training of deep networks
Alvarez, J. M. and M. Salzmann (2017) · 2017
Cited alongside, same era.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., S. Chen, and S. Pan (2017) · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Z. Liu, L. Maaten, and K. Weinberger (2017) · 2017
Optimization landscape and expressivity of deep CNNs
Nguyen, Q. and M. Hein (2018) · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Y. Bahri, J. Sohl-Dickstein, S. S. Schoenholz, and P. Pennington (2018) · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Arora, S., S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang (2019) · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S., X. Zhai, B. Poczos, and A. Singh (2019) · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and M. Carbin (2019) · 2019
Later among the works it cites.
On the impact of the activation function on deep neural networks training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep information propagation
Schoenholz, S., J. Gilmer, S. Ganguli, and J. Sohl-Dickstein (2017) · 2017
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
Yang, G. and S. Schoenholz (2017) · 2017
Cited alongside, same era.
Learning-compression algorithms for neural net pruning
Carreira-Perpiñán, M. and Y. Idelbayev (2018, June) · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., F. Gabriel, and C. Hongler (2018) · 2018
Cited alongside, same era.
Deep neural networks as Gaussian processes
Lee, J., Y. Bahri, R. Novak, S. Schoenholz, J. Pennington, and J. Sohl-Dickstein (2018) · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., T. Ajanthan, and P. H. Torr (2018) · 2018
Cited alongside, same era.
Hayou, S., A. Doucet, and J. Rousseau (2019) · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington (2019) · 2019
Later among the works it cites.
Metapruning: Meta learning for automatic neural network channel pruning
Liu, Z., H. Mu, X. Zhang, Z. Guo, X. Yang, K.-T. Cheng, and J. Sun (2019) · 2019
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Neyshabur, B., Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro (2019) · 2019
Later among the works it cites.
A mean field theory of batch normalization
Yang, G., J. Pennington, V. Rao, J. Sohl-Dickstein, and S. S. Schoenholz (2019) · 2019
Later among the works it cites.
A signal propagation perspective for pruning neural networks at initialization
Lee, N., T. Ajanthan, S. Gould, and P. H. S. Torr (2020) · 2020
Closest in time.
Group sparsity: The hinge between filter pruning and decomposition for network compression
Li, Y., S. Gu, C. Mayer, L. V. Gool, and R. Timofte (2020) · 2020
Closest in time.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., D. Kunin, D. L. Yamins, and S. Ganguli (2020) · 2020
Closest in time.
Picking winning tickets before training by preserving gradient flow
Wang, C., G. Zhang, and R. Grosse (2020) · 2020
Closest in time.
Stable resnet
Hayou, S., E. Clerico, B. He, G. Deligiannidis, A. Doucet, and J. Rousseau (2021) · 2021
Closest in time.