Fetching the paper…
Reading the bibliography…
To advance deep learning methodologies in the next decade, a theoretical framework for reasoning about modern neural networks is needed.
Schizophrenia: caused by a fault in programmed synaptic elimination during adolescence?
I. Feinberg · 1982
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
A. L. Blum and R. L. Rivest · 1992
Earlier work this paper cites.
The elements of statistical learning , volume 1
J. Friedman, T. Hastie, and R. Tibshirani · 2001
Earlier work this paper cites.
The organization of behavior: A neuropsychological theory
D. O. Hebb · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
N. Tishby and N. Zaslavsky · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2017
Earlier work this paper cites.
Dynamic routing between capsules
S. Sabour, N. Frosst, and G. E. Hinton · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
R. Shwartz-Ziv and N. Tishby · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
The local elasticity of neural networks
H. He and W. J. Su · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
S. Oymak and M. Soltanolkotabi · 2020
Later among the works it cites.
Prevalence of neural collapse during the terminal phase of deep learning training
V. Papyan, X. Han, and D. L. Donoho · 2020
Later among the works it cites.
Theoretical issues in deep networks
T. Poggio, A. Banburski, and Q. Liao · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
N. Razin and N. Cohen · 2020
Later among the works it cites.
On the generalization benefit of noise in stochastic gradient descent
S. Smith, E. Elsen, and S. De · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Wu, C. Ma, and W. E · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
A. Ilyas, S. Santurkar, L. Engstrom, B. Tran, and A. Madry · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Cited alongside, same era.
Uniform convergence may be unable to explain generalization in deep learning
V. Nagarajan and J. Z. Kolter · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain
Z.-Q. J. Xu, Y. Zhang, and Y. Xiao · 2019
Cited alongside, same era.
Understanding deep learning is also a job for physicists
L. Zdeborová · 2020
Later among the works it cites.
The modern mathematics of deep learning
J. Berner, P. Grohs, G. Kutyniok, and P. Petersen · 2021
Closest in time.
ReduNet: A white-box deep network from the principle of maximizing rate reduction
K. H. R. Chan, Y. Yu, C. You, H. Qi, J. Wright, and Y. Ma · 2021
Closest in time.
Toward better generalization bounds with locally elastic stability
Z. Deng, H. He, and W. J. Su · 2021
Closest in time.
The dawning of a new era in applied mathematics
W. E · 2021
Closest in time.
Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training
C. Fang, H. He, Q. Long, and W. J. Su · 2021
Closest in time.
How to represent part-whole hierarchies in a neural network
G. Hinton · 2021
Closest in time.
MlP-mixer: An all-MLP architecture for vision
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, D. Keysers, J. Uszkoreit, M. Lucic, and A. Dosovitskiy · 2021
Closest in time.
Noise or signal: The role of image backgrounds in object recognition
K. Xiao, L. Engstrom, A. Ilyas, and A. Madry · 2021
Closest in time.
Why over-parameterization of deep neural networks does not overfit?
Z.-H. Zhou · 2021
Closest in time.
Large models are parsimonious learners: Activation sparsity in trained transformers
Z. Li, C. You, S. Bhojanapalli, D. Li, A. S. Rawat, S. J. Reddi, K. Ye, F. Chern, F. Yu, R. Guo, and S. Kumar · 2022
Closest in time.
Moefication: Transformer feed-forward layers are mixtures of experts
Z. Zhang, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou · 2022
Closest in time.
A law of data separation in deep learning
H. He and W. J. Su · 2023
Closest in time.