Fetching the paper…
Reading the bibliography…
The last decade has seen blossoming research in deep learning theory attempting to answer, "Why does deep learning generalize?" A powerful shift in perspective precipitated this progress: the study of overparametrized models in the interpolation regime.
SGD on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 1905
Earlier work this paper cites.
Gradient-Based Neural DAG Learning
Lachapelle, S., Brouillard, P., Deleu, T., and Lacoste-Julien, S · 1906
Earlier work this paper cites.
Learning Neural Causal Models from Unknown Interventions
Ke, N. R., Bilaniuk, O., Goyal, A., Bauer, S., Larochelle, H., Schölkopf, B., Mozer, M. C., Pal, C., and Bengio, Y · 1910
Earlier work this paper cites.
Learning under Model Misspecification: Applications to Variational and Ensemble methods
Masegosa, A. R · 1912
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R · 1964
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. and Chervonenkis, A. Y · 1971
Earlier work this paper cites.
Finite State Automata and Simple Recurrent Networks
Cleeremans, A., Servan-Schreiber, D., and McClelland, J. L · 1989
Earlier work this paper cites.
Independent component analysis, a new concept?
Comon, P · 1994
Earlier work this paper cites.
The language instinct
Pinker, S · 1994
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Schmidhuber, J · 1997
Earlier work this paper cites.
On prediction by data compression
Vitányi, P. and Li, M · 1997
Earlier work this paper cites.
On tables of random numbers
Kolmogorov, A · 1998
Earlier work this paper cites.
Nonlinear independent component analysis: Existence and uniqueness results
Hyvärinen, A. and Pajunen, P · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity, 2000
Hutter, M · 2000
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vapnik, V · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
On Layer Normalization in the Transformer Architecture, June 2020
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2002
Earlier work this paper cites.
A Survey of Neural Networks and Formal Languages, June 2020
Ackerman, J. and Cybenko, G · 2006
Earlier work this paper cites.
Object-Centric Learning with Slot Attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2006
Earlier work this paper cites.
A Linear Non-Gaussian Acyclic Model for Causal Discovery
Shimizu, S., Hoyer, P. O., Hyvarinen, A., and Kerminen, A · 2006
Earlier work this paper cites.
Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse Coding
Klindt, D., Schott, L., Sharma, Y., Ustyuzhaninov, I., Brendel, W., Bethge, M., and Paiton, D · 2007
Earlier work this paper cites.
Nonlinear causal discovery with additive noise models
Hoyer, P., Janzing, D., Mooij, J. M., Peters, J., and Schölkopf, B · 2008
Earlier work this paper cites.
Systematic generalization on gscan with language conditioned embedding
Gao, T., Huang, Q., and Mooney, R. J · 2009
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Pearl, J · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Independent component analysis: recent advances
Hyvärinen, A · 2011
Earlier work this paper cites.
Prediction of time series by statistical learning: general losses and fast rates, 2012
Alquier, P., Li, X., and Wintenberger, O · 2012
Earlier work this paper cites.
A survey of l1 regression
Vidaurre, D., Bielza, C., and Larrañaga, P · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
A new pac-bayesian perspective on domain adaptation
Germain, P., Habrard, A., Laviolette, F., and Morvant, E · 2016
Earlier work this paper cites.
Unsupervised Feature Extraction by Time-Contrastive Learning and Nonlinear ICA
Hyvarinen, A. and Morioka, H · 2016
Earlier work this paper cites.
Controlling bias in adaptive data analysis using information theory
Russo, D. and Zou, J · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Simpler pac-bayesian bounds for hostile data
Alquier, P. and Guedj, B · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks, 2017
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., and Lacoste-Julien, S · 2017
Earlier work this paper cites.
Sharp Minima Can Generalize For Deep Nets, May 2017
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization, 2017
Gunasekar, S., Woodworth, B., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Earlier work this paper cites.
Is Maximum Likelihood Useful for Representation Learning?
Huszár, F · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Earlier work this paper cites.
Input–output maps are strongly biased towards simple outputs
Dingle, K., Camargo, C., and Louis, A · 2018
Earlier work this paper cites.
Generalisation in humans and deep neural networks
Geirhos, R., Temme, C. R. M., Rauber, J., Schütt, H. H., Bethge, M., and Wichmann, F. A · 2018
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization, 2019
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Cited alongside, same era.
Systematic generalization: What is required and can it be learned?
Bahdanau, D., Murty, S., Noukhovitch, M., Nguyen, T. H., de Vries, H., and Courville, A · 2019
Cited alongside, same era.
Random deep neural networks are biased towards simple functions, 2019
De Palma, G., Kiani, B. T., and Lloyd, S · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
PyTorch Lightning, March 2019
Falcon, W. and The PyTorch Lightning team · 2019
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks, 2019
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2019
Inductive Biases for Object-Centric Representations in the Presence of Complex Textures, August 2022
Papa, S., Winther, O., and Dittadi, A · 2022
Later among the works it cites.
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, January 2022
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Later among the works it cites.
A general framework for pac-bayes bounds for meta-learning, 06 2022
Rezazadeh, A · 2022
Later among the works it cites.
Improving systematic generalization through modularity and augmentation, 2022
Ruis, L. and Lake, B · 2022
Later among the works it cites.
Content suppresses style: dimensionality collapse in contrastive learning
Rusak, E., Reizinger, P., Zimmermann, R. S., Bringmann, O., and Brendel, W · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fantastic generalization measures and where to find them, 2019
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Cited alongside, same era.
Compositional generalization for primitive substitutions
Li, Y., Zhao, L., Wang, J., and Hestness, J · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization, January 2019
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
On the spectral bias of neural networks, 2019
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F. A., Bengio, Y., and Courville, A · 2019
Cited alongside, same era.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Perez, G., Camargo, C. Q., and Louis, A. A · 2019
Cited alongside, same era.
Understanding Contrastive Learning Requires Incorporating Inductive Biases, February 2022
Saunshi, N., Ash, J., Goel, S., Misra, D., Zhang, C., Arora, S., Kakade, S., and Krishnamurthy, A · 2022
Later among the works it cites.
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers, January 2022
Tay, Y., Dehghani, M., Rao, J., Fedus, W., Abnar, S., Chung, H. W., Narang, S., Yogatama, D., Vaswani, A., and Metzler, D · 2022
Later among the works it cites.
Tracking the contribution of inductive bias to individualised internal models
Török, B., Nagy, D. G., Kiss, M., Janacsek, K., Németh, D., and Orbán, G · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Later among the works it cites.
An Explanation of In-context Learning as Implicit Bayesian Inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Later among the works it cites.
On Generalization of Adversarial Imitation Learning and Beyond, February 2022
Xu, T., Li, Z., Yu, Y., and Luo, Z.-Q · 2022
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Later among the works it cites.
Llemma: An open language model for mathematics, 2023
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Later among the works it cites.
Predicting Ordinary Differential Equations with Transformers
Becker, S., Klein, M., Neitz, A., Parascandolo, G., and Kilbertus, N · 2023
Later among the works it cites.
Provably Learning Object-Centric Representations, May 2023
Brady, J., Zimmermann, R. S., Sharma, Y., Schölkopf, B., von Kügelgen, J., and Brendel, W · 2023
Later among the works it cites.
Language modeling is compression, 2023
Delétang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J · 2023
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality, 2023
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y · 2023
Later among the works it cites.
Reinforcement Learning for Language Models, 2023
Goldberg, Y · 2023
Later among the works it cites.
The no free lunch theorem, kolmogorov complexity, and the role of inductive biases in machine learning, 2023
Goldblum, M., Finzi, M., Rowan, K., and Wilson, A. G · 2023
Later among the works it cites.
Hyvärinen, A., Khemakhem, I., and Monti, R · 2023
Later among the works it cites.
Mathprompter: Mathematical reasoning using large language models, 2023
Imani, S., Du, L., and Shrivastava, H · 2023
Later among the works it cites.
Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation, July 2023
Lachapelle, S., Mahajan, D., Mitliagkas, I., and Lacoste-Julien, S · 2023
Later among the works it cites.
Human-like systematic generalization through a meta-learning neural network
Lake, B. M. and Baroni, M · 2023
Later among the works it cites.
Non-vacuous generalization bounds for large language models
Lotfi, S., Finzi, M., Kuang, Y., Rudner, T., Goldblum, M., and Wilson, A · 2023
Later among the works it cites.
Rotating Features for Object Discovery, June 2023
Löwe, S., Lippe, P., Locatello, F., and Welling, M · 2023
Later among the works it cites.
Understanding generalization in the interpolation regime using the rate function, 2023
Masegosa, A. R. and Ortega, L. A · 2023
Later among the works it cites.
Pac-bayesian generalization bounds for adversarial generative models
Mbacke, S. D., Clerc, F., and Germain, P · 2023
Later among the works it cites.
Formal languages and neural models for learning on sequences
Merrill, W · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI et al · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Later among the works it cites.
Neural networks trained with sgd learn distributions of increasing complexity, 2023
Refinetti, M., Ingrosso, A., and Goldt, S · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., Kohli, P., and Fawzi, A · 2023
Later among the works it cites.
Testing the general deductive reasoning capacity of large language models using ood examples, 2023
Saparov, A., Pang, R. Y., Padmakumar, V., Joshi, N., Kazemi, S. M., Kim, N., and He, H · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent, May 2023
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Trained Transformers Learn Linear Models In-Context, October 2023
Zhang, R., Frei, S., and Bartlett, P. L · 2023
Later among the works it cites.
On Provable Length and Compositional Generalization, February 2024
Ahuja, K. and Mansouri, A · 2024
Closest in time.
Learning universal predictors, 2024
Grau-Moya, J., Genewein, T., Hutter, M., Orseau, L., Delétang, G., Catt, E., Ruoss, A., Wenliang, L. K., Mattern, C., Aitchison, M., and Veness, J · 2024
Closest in time.
Han, S. and Padó, S · 2024
Closest in time.
Linearity of Relation Decoding in Transformer Language Models, February 2024
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D · 2024
Closest in time.
Understanding the Effects of RLHF on LLM Generalisation and Diversity, February 2024
Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R · 2024
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Mahankali, A. V., Hashimoto, T., and Ma, T · 2024
Closest in time.
Language Models Implement Simple Word2Vec-style Vector Arithmetic, April 2024
Merullo, J., Eickhoff, C., and Pavlick, E · 2024
Closest in time.
Ramesh, R., Lubana, E. S., Khona, M., Dick, R. P., and Tanaka, H · 2024
Closest in time.