Fetching the paper…
Reading the bibliography…
The ability of deep neural networks to generalise well even when they interpolate their training data has been explained using various "simplicity biases".
Characteristic Functions
Lukacs, E · 1970
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
Le Cun, Y., Kanter, I., and Solla, S. A · 1991
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
Krogh, A. and Hertz, J. A · 1992
Earlier work this paper cites.
Generalization in a large committee machine
Schwarze, H. and Hertz, J · 1992
Earlier work this paper cites.
Exact Solution for On-Line Learning in Multilayer Neural Networks
Saad, D. and Solla, S · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Statistical Mechanics of Learning
Engel, A. and Van den Broeck, C · 2001
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Li, F · 2009
Earlier work this paper cites.
Tensor decompositions and applications
Kolda, T. G. and Bader, B. W · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep learning and hierarchical generative models
Mossel, E · 2016
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2017
Earlier work this paper cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A. C., Bengio, Y., and Lacoste-Julien, S · 2017
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
SGDR: stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
SVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2018
Cited alongside, same era.
A spectral approach to generalization and optimization in neural networks
Farnia, F., Zhang, J., and Tse, D · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Cited alongside, same era.
On the spectrum of random features maps of high dimensional data
Liao, Z. and Couillet, R · 2018
Cited alongside, same era.
Tensor methods in statistics
McCullagh, P · 2018
Directional convergence and alignment in deep learning
Ji, Z. and Telgarsky, M · 2020
Later among the works it cites.
Implicit bias of gradient descent for mean squared error regression with wide neural networks
Jin, H. and Montúfar, G · 2020
Later among the works it cites.
Pre-training without natural images
Kataoka, H., Okayasu, K., Matsumoto, A., Yamagata, E., Yamada, R., Inoue, N., Nakamura, A., and Satoh, Y · 2020
Later among the works it cites.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Later among the works it cites.
Bad global minima exist and SGD can reach them
Liu, S., Papailiopoulos, D. S., and Achlioptas, D · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., and Srebro, N · 2018
Cited alongside, same era.
Understanding training and generalization in deep learning by fourier analysis
Xu, Z. J · 2018
Cited alongside, same era.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2019
Cited alongside, same era.
Intrinsic dimension of data representations in deep neural networks
Ansuini, A., Laio, A., Macke, J. H., and Zoccolan, D · 2019
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Cited alongside, same era.
Later among the works it cites.
What do neural networks learn when trained with random labels?
Maennel, H., Alabdulmohsin, I. M., Tolstikhin, I. O., Baldock, R. J. N., Bousquet, O., Gelly, S., and Keysers, D · 2020
Later among the works it cites.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
Mignacco, F., Krzakala, F., Urbani, P., and Zdeborová, L · 2020
Later among the works it cites.
Distributional generalization: A new kind of generalization
Nakkiran, P. and Bansal, Y · 2020
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Spigler, S., Geiger, M., and Wyart, M · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Does pretraining for summarization require knowledge transfer?
Krishna, K., Bigham, J., and Lipton, Z. C · 2021
Later among the works it cites.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Loureiro, B., Gerbelot, C., Cui, H., Goldt, S., Krzakala, F., Mézard, M., and Zdeborová, L · 2021
Later among the works it cites.
Learning gaussian mixtures with generalized linear models: Precise asymptotics in high-dimensions
Loureiro, B., Sicuro, G., Gerbelot, C., Pacco, A., Krzakala, F., and Zdeborová, L · 2021
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Mei, S. and Montanari, A · 2021
Later among the works it cites.
The bootstrap framework: Generalization through the lens of online optimization
Nakkiran, P., Neyshabur, B., and Sedghi, H · 2021
Later among the works it cites.
The intrinsic dimension of images and its impact on learning
Pope, P., Zhu, C., Abdelkader, A., Goldblum, M., and Goldstein, T · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Refinetti, M., Goldt, S., Krzakala, F., and Zdeborová, L · 2021
Later among the works it cites.
Minimal CIFAR10
van Amersfoort, J · 2021
Later among the works it cites.
Implicit bias of MSE gradient optimization in underparameterized neural networks
Bowman, B. and Montúfar, G · 2022
Closest in time.
Decomposing neural networks as mappings of correlation functions
Fischer, K., René, A., Keup, C., Layer, M., Dahmen, D., and Helias, M · 2022
Closest in time.
Gaussian universality of linear classifiers with random labels in high-dimension
Gerace, F., Krzakala, F., Loureiro, B., Stephan, L., and Zdeborová, L · 2022
Closest in time.
The gaussian equivalence of generative models for learning with shallow neural networks
Goldt, S., Loureiro, B., Reeves, G., Krzakala, F., Mezard, M., and Zdeborová, L · 2022
Closest in time.
Universality laws for high-dimensional learning with random features
Hu, H. and Lu, Y. M · 2022
Closest in time.
Data-driven emergence of convolutional structure in neural networks
Ingrosso, A. and Goldt, S · 2022
Closest in time.
The dynamics of representation learning in shallow, non-linear autoencoders
Refinetti, M. and Goldt, S · 2022
Closest in time.
On the implicit bias in deep-learning algorithms
Vardi, G · 2022
Closest in time.
Overcoming the spectral bias of neural value approximation
Yang, G., Ajay, A., and Agrawal, P · 2022
Closest in time.