Fetching the paper…
Reading the bibliography…
The remarkable performance of overparameterized deep neural networks (DNNs) must arise from an interplay between network architecture, training algorithms, and structure in the data.
On the method of theoretical physics
Einstein, A · 1934
Earlier work this paper cites.
A preliminary report on a general theory of inductive inference (Citeseer, 1960)
Solomonoff, R. J · 1960
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R. J · 1964
Earlier work this paper cites.
Three approaches to the quantitative definition of information
Kolmogorov, A. N · 1968
Earlier work this paper cites.
On the simplicity and speed of programs for computing infinite sets of natural numbers
Chaitin, G. J · 1969
Earlier work this paper cites.
Perceptrons: an introduction to computational geometry (MIT press, 1969)
Minsky, M. & Papert, S. A · 1969
Earlier work this paper cites.
Laws of information conservation (nongrowth) and aspects of the foundation of probability theory
Levin, L · 1974
Earlier work this paper cites.
On the complexity of finite sequences
Lempel, A. & Ziv, J · 1976
Earlier work this paper cites.
On the connection between the complexity and credibility of inferred models
Pearl, J · 1978
Earlier work this paper cites.
Modeling by shortest data description
Rissanen, J · 1978
Earlier work this paper cites.
The need for biases in learning generalizations (rutgers computer science tech. rept. cbm-tr-117)
Mitchell, T. M · 1980
Earlier work this paper cites.
Occam’s razor
Blumer, A., Ehrenfeucht, A., Haussler, D. & Warmuth, M. K · 1987
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Logik der Forschung (JCB Mohr Tübingen, 1989)
Popper, K. R · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K · 1991
Earlier work this paper cites.
Overfitting avoidance as bias
Schaffer, C · 1993
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Neal, R. M · 1994
Earlier work this paper cites.
A conservation law for generalization performance
Schaffer, C · 1994
Earlier work this paper cites.
Reflections after refereeing papers for nips
Breiman, L · 1995
Earlier work this paper cites.
Are efficient deep representations learnable?
Nye, M. & Saxe, A · 1995
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
Wolpert, D. H · 1996
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Kohavi, R., Wolpert, D. H. et al · 1996
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Schmidhuber, J · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. & Macready, W. G · 1997
Earlier work this paper cites.
The miraculous universal distribution
Kirchherr, W., Li, M. & Vitányi, P · 1997
Earlier work this paper cites.
Some pac-bayesian theorems
McAllester, D. A · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
McAllester, D. A · 1999
Earlier work this paper cites.
The role of occam’s razor in knowledge discovery
Domingos, P · 1999
Earlier work this paper cites.
Pac-bayesian model averaging
McAllester, D. A · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Hutter, M · 2000
Earlier work this paper cites.
A unified bias-variance decomposition
Domingos, P · 2000
Earlier work this paper cites.
The speed prior: a new simplicity measure yielding near-optimal computable predictions
Schmidhuber, J · 2002
Earlier work this paper cites.
Information theory, inference and learning algorithms (Cambridge university press, 2003)
MacKay, D. J · 2003
Earlier work this paper cites.
A meeting with enrico fermi
Dyson, F · 2004
Earlier work this paper cites.
Simplicity
Baker, A · 2004
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability (Springer Science & Business Media, 2004)
Hutter, M · 2004
Earlier work this paper cites.
Estimating the entropy rate of spike trains via lempel-ziv complexity
Amigó, J. M., Szczepański, J., Wajnryb, E. & Sanchez-Vives, M. V · 2004
Earlier work this paper cites.
Intrinsic dimensionality estimation of submanifolds in rd
Hein, M. & Audibert, J.-Y · 2005
Earlier work this paper cites.
A new solution to the puzzle of simplicity
Kelly, K. T · 2007
Earlier work this paper cites.
Reliable Reasoning: Induction and Statistical Learning Theory (MIT Press, 2007)
Harman, G. & Kulkarni, S · 2007
Earlier work this paper cites.
The minimum description length principle (MIT Press, 2007)
Grünwald, P. D · 2007
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications (Springer-Verlag New York Inc, 2008)
Li, M. & Vitanyi, P · 2008
Earlier work this paper cites.
Bias vs variance decomposition for regression and classification
Geurts, P · 2009
Earlier work this paper cites.
Entropy estimation of very short symbolic sequences
Lesne, A., Blanc, J.-L. & Pezard, L · 2009
Earlier work this paper cites.
A philosophical treatise of universal induction
Rathmanner, S. & Hutter, M · 2011
Earlier work this paper cites.
No free lunch versus occam’s razor in supervised learning
Lattimore, T. & Hutter, M · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms (Cambridge university press, 2014)
Shalev-Shwartz, S. & Ben-David, S · 2014
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y. & Hinton, G · 2015
Earlier work this paper cites.
Ockham’s razors (Cambridge University Press, 2015)
Sober, E · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B. & Vinyals, O · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J. & Ganguli, S · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M. & Tang, P. T. P · 2016
Cited alongside, same era.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H. & Graf, H. P · 2016
Cited alongside, same era.
Finite versus infinite neural networks: an empirical study
Lee, J. et al · 2020
Later among the works it cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Bordelon, B., Canatar, A. & Pehlevan, C · 2020
Later among the works it cites.
Generalization bounds for deep learning
Valle-Pérez, G. & Louis, A. A · 2020
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Spigler, S., Geiger, M. & Wyart, M · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P. & Netrapalli, P · 2020
Later among the works it cites.
A brief prehistory of double descent
Loog, M., Viering, T., Mey, A., Krijthe, J. H. & Tax, D. M · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Relating and contrasting plain and prefix kolmogorov complexity
Bauwens, B · 2016
Cited alongside, same era.
Solomonoff prediction and occam’s razor
Sterkenburg, T. F · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D. et al · 2017
Cited alongside, same era.
Why does deep and cheap learning work so well?
Lin, H. W., Tegmark, M. & Rolnick, D · 2017
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S. & Sohl-Dickstein, J · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P. et al · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I. & Soudry, D · 2017
Cited alongside, same era.
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S., Saxe, A. M. & Sompolinsky, H · 2020
Later among the works it cites.
Feature learning in infinite-width neural networks
Yang, G. & Hu, E. J · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Geiger, M., Spigler, S., Jacot, A. & Wyart, M · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Wilson, A. G. & Izmailov, P · 2020
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G. & Tsigler, A · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Geirhos, R. et al · 2020
Later among the works it cites.
Generic predictions of output probability based on complexities of inputs and outputs
Dingle, K., Pérez, G. V. & Louis, A. A · 2020
Later among the works it cites.
Boolean threshold networks as models of genotype-phenotype maps
Camargo, C. Q. & Louis, A. A · 2020
Later among the works it cites.
Space of functions computed by deep-layered machines
Mozeika, A., Li, B. & Saad, D · 2020
Later among the works it cites.
Belkin, M · 2021
Later among the works it cites.
Learning curves for overparametrized deep neural networks: A field theory perspective
Cohen, O., Malka, O. & Ringel, Z · 2021
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Canatar, A., Bordelon, B. & Pehlevan, C · 2021
Later among the works it cites.
Double-descent curves in neural networks: a new perspective using gaussian processes
Harzli, O. E., Valle-Pérez, G. & Louis, A. A · 2021
Later among the works it cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Cui, H., Loureiro, B., Krzakala, F. & Zdeborová, L · 2021
Later among the works it cites.
Neural tangent kernel eigenvalues accurately predict generalization
Simon, J. B., Dickens, M. & DeWeese, M. R · 2021
Later among the works it cites.
The low-rank simplicity bias in deep networks
Huh, M. et al · 2021
Later among the works it cites.
Is sgd a bayesian sampler? well, almost
Mingard, C., Valle-Pérez, G., Skalse, J. & Louis, A. A · 2021
Later among the works it cites.
Physics-informed machine learning
Karniadakis, G. E. et al · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P. et al · 2021
Later among the works it cites.
Feature learning and signal propagation in deep neural networks (2021)
Lou, Y., Mingard, C., Nam, Y. & Hayou, S · 2021
Later among the works it cites.
Implicit regularization via neural feature alignment
Baratin, A. et al · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Geiping, J., Goldblum, M., Pope, P. E., Moeller, M. & Goldstein, T · 2021
Later among the works it cites.
Kernel interpolation as a bayes point machine
Bernstein, J., Farhang, A. & Yue, Y · 2021
Later among the works it cites.
Simplicity bias in transformers and their ability to learn sparse boolean functions
Bhattamishra, S., Patel, A., Kanade, V. & Blunsom, P · 2022
Later among the works it cites.
Investigating generalization by controlling normalized margin
Farhang, A. R., Bernstein, J. D., Tirumala, K., Liu, Y. & Yue, Y · 2022
Later among the works it cites.
Biological convolutions improve dnn robustness to noise and generalisation
Evans, B. D., Malhotra, G. & Bowers, J. S · 2022
Later among the works it cites.
The principles of deep learning theory (Cambridge University Press, 2022)
Roberts, D. A., Yaida, S. & Hanin, B · 2022
Later among the works it cites.
Benign, tempered, or catastrophic: A taxonomy of overfitting
Mallinar, N. et al · 2022
Later among the works it cites.
Neural networks trained with sgd learn distributions of increasing complexity
Refinetti, M., Ingrosso, A. & Goldt, S · 2022
Later among the works it cites.
Overview frequency principle/spectral bias in deep learning
Xu, Z.-Q. J., Zhang, Y. & Luo, T · 2022
Later among the works it cites.
Symmetry and simplicity spontaneously emerge from the algorithmic nature of evolution
Johnston, I. G. et al · 2022
Later among the works it cites.
Simple models in complex worlds: Occam’s razor and statistical learning theory
Bargagli Stoffi, F. J., Cevolani, G. & Gnecco, G · 2022
Later among the works it cites.
A dilemma for solomonoff prediction
Neth, S · 2022
Later among the works it cites.
Loss landscapes are all you need: Neural network generalization can be explained without the implicit bias of gradient descent
Chiang, P.-y. et al · 2023
Closest in time.
Language modeling is compression
Delétang, G. et al · 2023
Closest in time.
Separation of scales and a thermodynamic description of feature learning in some cnns
Seroussi, I., Naveh, G. & Ringel, Z · 2023
Closest in time.
A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit
Pacelli, R., Ariosto, S. & Ginelli, F · 2023
Closest in time.
A spectral condition for feature learning
Yang, G., Simon, J. B. & Bernstein, J · 2023
Closest in time.
Robustness and stability of spin-glass ground states to perturbed interactions
Mohanty, V. & Louis, A. A · 2023
Closest in time.
Grau-Moya, J. et al · 2024
Closest in time.
Li, Z. et al · 2024
Closest in time.