Fetching the paper…
Reading the bibliography…
No free lunch theorems for supervised learning state that no learner can solve all problems or that all learners achieve exactly the same accuracy on average over a uniform distribution on learning problems.
On tables of random numbers
Kolmogorov, A. N · 1963
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R. J · 1964
Earlier work this paper cites.
Information-theoretic limitations of formal systems
Chaitin, G. J · 1974
Earlier work this paper cites.
Arithmetic coding for data compression
Witten, I. H., Neal, R. M., and Cleary, J. G · 1987
Earlier work this paper cites.
Chaitin-kolmogorov complexity and generalization in neural networks
Pearlmutter, B. and Rosenfeld, R · 1990
Earlier work this paper cites.
Complexity regularization with application to artificial neural networks
Barron, A. R · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, G. E. and Van Camp, D · 1993
Earlier work this paper cites.
A conservation law for generalization performance
Schaffer, C · 1994
Earlier work this paper cites.
Kolmogorov complexity of finite sequences and recognition of different preictal eeg patterns
Petrosian, A · 1995
Earlier work this paper cites.
For every generalization action, is there really an equal and opposite reaction? analysis of the conservation law for generalization performance
Rao, R. B., Gordon, D., and Spears, W · 1995
Earlier work this paper cites.
No free lunch theorems for search
Wolpert, D. H., Macready, W. G., et al · 1995
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M · 1996
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
Wolpert, D. H · 1996
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Schmidhuber, J · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. and Macready, W. G · 1997
Earlier work this paper cites.
Some PAC-Bayesian theorems
McAllester, D. A · 1998
Earlier work this paper cites.
Adaptive and learning systems for signal processing communications, and control
Vapnik, V. N · 1998
Earlier work this paper cites.
Algorithm performance and problem structure for flow-shop scheduling
Watson, J.-P., Barbulescu, L., Howe, A. E., and Whitley, L. D · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Hutter, M · 2000
Earlier work this paper cites.
Kolmogorov complexity
Fortnow, L · 2001
Earlier work this paper cites.
Bounds for averaging classifiers
Langford, J. and Seeger, M · 2001
Earlier work this paper cites.
Simple explanation of the no-free-lunch theorem and its implications
Ho, Y.-C. and Pepyne, D. L · 2002
Earlier work this paper cites.
A perspective view and survey of meta-learning
Vilalta, R. and Drissi, Y · 2002
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I · 2003
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Dalal, N. and Triggs, B · 2005
Earlier work this paper cites.
Toward a justification of meta-learning: Is the no free lunch theorem a show-stopper
Giraud-Carrier, C. and Provost, F · 2005
Earlier work this paper cites.
Complexity theory and the no free lunch theorem
Whitley, D. and Watson, J. P · 2005
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Mnih, A. and Hinton, G · 2007
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications , volume 3
Li, M., Vitányi, P., et al · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, A · 2009
Earlier work this paper cites.
Rademacher complexity bounds for non-iid processes
Mohri, M. and Rostamizadeh, A · 2009
Cited alongside, same era.
A philosophical treatise of universal induction
Rathmanner, S. and Hutter, M · 2011
Cited alongside, same era.
“no free lunch” theorems applied to the calibration of traffic simulation models
Ciuffo, B. and Punzo, V · 2013
Cited alongside, same era.
No free lunch versus occam’s razor in supervised learning
Lattimore, T. and Hutter, M · 2013
Cited alongside, same era.
Hidden factors and hidden topics: understanding rating dimensions with review text
McAuley, J. and Leskovec, J · 2013
Cited alongside, same era.
Gaussian process kernels for pattern discovery and extrapolation
Wilson, A. and Adams, R · 2013
Cited alongside, same era.
Optimal regularization can mitigate double descent
Nakkiran, P., Venkat, P., Kakade, S. M., and Ma, T · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Wilson, A. G. and Izmailov, P · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as sgd
Wu, J., Hu, W., Xiong, H., Huan, J., Braverman, V., and Zhu, Z · 2020
Later among the works it cites.
The information complexity of learning tasks, their structure and their distance
Achille, A., Paolini, G., Mbeng, G., and Soatto, S · 2021
Later among the works it cites.
Sgd generalizes better than gd (and regularization doesn’t help)
Amir, I., Koren, T., and Livni, R · 2021
Later among the works it cites.
A close look at deep learning with small data
Brigato, L. and Iocchi, L · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do we need hundreds of classifiers to solve real world classification problems?
Fernández-Delgado, M., Cernadas, E., Barro, S., and Amorim, D · 2014
Cited alongside, same era.
Automatic construction and natural-language description of nonparametric regression models
Lloyd, J., Duvenaud, D., Grosse, R., Tenenbaum, J., and Ghahramani, Z · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Cited alongside, same era.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Zenil, H., Marshall, J. A., and Tegnér, J · 2015
Cited alongside, same era.
On the role of data in pac-bayes bounds
Dziugaite, G. K., Hsu, K., Gharbieh, W., Arpino, G., and Roy, D · 2021
Later among the works it cites.
A practical method for constructing equivariant multilayer perceptrons for arbitrary matrix groups
Finzi, M., Welling, M., and Wilson, A. G · 2021
Later among the works it cites.
A theoretical analysis of the repetition problem in text generation
Fu, Z., Lam, W., So, A. M.-C., and Shi, B · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Geiping, J., Goldblum, M., Pope, P. E., Moeller, M., and Goldstein, T · 2021
Later among the works it cites.
What are bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G. G · 2021
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., et al · 2021
Later among the works it cites.
Vision transformer for small-size datasets
Lee, S. H., Lee, S., and Song, B. C · 2021
Later among the works it cites.
Limitations of autoregressive models and their alternatives
Lin, C.-C., Jaech, A., Li, X., Gormley, M. R., and Eisner, J · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Turing-universal learners with optimal scaling laws
Nakkiran, P · 2021
Later among the works it cites.
Tighter risk certificates for neural networks
Pérez-Ortiz, M., Rivasplata, O., Shawe-Taylor, J., and Szepesvári, C · 2021
Later among the works it cites.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Schwarzschild, A., Borgnia, E., Gupta, A., Huang, F., Vishkin, U., Goldblum, M., and Goldstein, T · 2021
Later among the works it cites.
Saint: Improved neural networks for tabular data via row attention and contrastive pre-training
Somepalli, G., Goldblum, M., Schwarzschild, A., Bruss, C. B., and Goldstein, T · 2021
Later among the works it cites.
End-to-end algorithm synthesis with recurrent networks: Extrapolation without overthinking
Bansal, A., Schwarzschild, A., Borgnia, E., Emam, Z., Huang, F., Goldblum, M., and Goldstein, T · 2022
Later among the works it cites.
Deep symbolic regression for recurrent sequences
d’Ascoli, S., Kamienny, P.-A., Lample, G., and Charton, F · 2022
Later among the works it cites.
Why do tree-based models still outperform deep learning on tabular data?
Grinsztajn, L., Oyallon, E., and Varoquaux, G · 2022
Later among the works it cites.
Tabpfn: A transformer that solves small tabular classification problems in a second
Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F · 2022
Later among the works it cites.
The low-rank simplicity bias in deep networks
Huh, M., Mobahi, H., Zhang, R., Cheung, B., Agrawal, P., and Isola, P · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2022
Later among the works it cites.
PAC-Bayes compression bounds so tight that they can explain generalization
Lotfi, S., Finzi, M., Kapoor, S., Potapczynski, A., Goldblum, M., and Wilson, A. G · 2022
Later among the works it cites.
Transformers can do bayesian inference
Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F · 2022
Later among the works it cites.
Trockman, A. and Kolter, J. Z · 2022
Later among the works it cites.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., and Goodman, N. D · 2022
Later among the works it cites.
Loss landscapes are all you need: Neural network generalization can be explained without the implicit bias of gradient descent
Chiang, P.-y., Ni, R., Miller, D. Y., Bansal, A., Geiping, J., Goldblum, M., and Goldstein, T · 2023
Closest in time.
The lie derivative for measuring learned equivariance
Gruver, N., Finzi, M., Goldblum, M., and Wilson, A. G · 2023
Closest in time.
Turing complete transformers: Two transformers are more powerful than one
Upadhyay, S. K. and Ginsberg, E. J · 2023
Closest in time.
Transformer-based models are not yet perfect at learning to emulate structural recursion
Zhang, D., Tigges, C., Zhang, Z., Biderman, S., Raginsky, M., and Ringer, T · 2024
Closest in time.