Fetching the paper…
Reading the bibliography…
The distributional simplicity bias (DSB) posits that neural networks learn low-order moments of the data distribution first, before moving on to higher-order correlations.
Frequency principle: Fourier analysis sheds light on deep neural networks
Xu, Z.-Q. J., Zhang, Y., Luo, T., Xiao, Y., and Ma, Z · 1901
Earlier work this paper cites.
Information theory and statistical mechanics
Jaynes, E. T · 1957
Earlier work this paper cites.
Fonctions de répartition à n dimensions et leurs marges
Sklar, M · 1959
Earlier work this paper cites.
Maximum-entropy distributions having prescribed first and second moments (corresp.)
Dowson, D. and Wragg, A · 1973
Earlier work this paper cites.
The fréchet distance between multivariate normal distributions
Dowson, D. and Landau, B · 1982
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
Kozachenko, L. F. and Leonenko, N. N · 1987
Earlier work this paper cites.
Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions
Vallet, F., Cailton, J.-G., and Refregier, P · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Probability distributions and maximum entropy
Conrad, K · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
On the strong convergence of the optimal linear shrinkage estimator for large dimensional covariance matrix
Bodnar, T., Gupta, A. K., and Parolya, N · 2014
Earlier work this paper cites.
Optimal transport for applied mathematicians
Santambrogio, F · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Typical sets and the curse of dimensionality
Carpenter, B · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Cited alongside, same era.
Spreading vectors for similarity search
Sablayrolles, A., Douze, M., Schmid, C., and Jégou, H · 2018
Cited alongside, same era.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Later among the works it cites.
Implicit regularization via neural feature alignment
Baratin, A., George, T., Laurent, C., Hjelm, R. D., Lajoie, G., Vincent, P., and Lacoste-Julien, S · 2021
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Canatar, A., Bordelon, B., and Pehlevan, C · 2021
Later among the works it cites.
On the adequacy of untuned warmup for adaptive optimization
Ma, J. and Yarats, D · 2021
Later among the works it cites.
Deep frequency principle towards understanding why deeper learning is faster
Xu, Z. J. and Zhou, H · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Valle-Perez, G., Camargo, C. Q., and Louis, A. A · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Cited alongside, same era.
On the inductive bias of neural tangent kernels
Bietti, A. and Mairal, J · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., and Courville, A · 2019
Cited alongside, same era.
Frequency bias in neural networks for input of non-uniform density
Basri, R., Galun, M., Geifman, A., Jacobs, D., Kasten, Y., and Kritchman, S · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Later among the works it cites.
The grammar-learning trajectories of neural language models
Choshen, L., Hacohen, G., Weinshall, D., and Abend, O · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Later among the works it cites.
Loss landscapes are all you need: Neural network generalization can be explained without the implicit bias of gradient descent
Chiang, P., Ni, R., Miller, D. Y., Bansal, A., Geiping, J., Goldblum, M., and Goldstein, T · 2023
Later among the works it cites.
Battle of the backbones: A large-scale comparison of pretrained models across computer vision tasks
Goldblum, M., Souri, H., Ni, R., Shu, M., Prabhu, V. U., Somepalli, G., Chattopadhyay, P., Ibrahim, M., Bardes, A., Hoffman, J., et al · 2023
Later among the works it cites.
Neural networks trained with sgd learn distributions of increasing complexity
Refinetti, M., Ingrosso, A., and Goldt, S · 2023
Later among the works it cites.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., and Xie, S · 2023
Later among the works it cites.