Fetching the paper…
Reading the bibliography…
A central goal of machine learning is generalization.
Meta-learning of sequential strategies
Ortega, P. A., Wang, J. X., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., Sprechmann, P., et al · 1905
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R. J · 1964
Earlier work this paper cites.
Three approaches to the quantitative definition of information
Kolmogorov, A. N · 1965
Earlier work this paper cites.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J. M., Bengio, S., and Nakkiran, P · 1965
Earlier work this paper cites.
On the length of programs for computing finite binary sequences
Chaitin, G. J · 1966
Earlier work this paper cites.
Arithmetic coding for data compression
Witten, I. H., Neal, R. M., and Cleary, J. G · 1987
Earlier work this paper cites.
Statistical learning theory
Vapnik, V. N., Vapnik, V., et al · 1998
Earlier work this paper cites.
Kolmogorov complexity
Fortnow, L · 2000
Earlier work this paper cites.
Learning to Learn Using Gradient Descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 2001
Earlier work this paper cites.
Kolmogorov complexity and information theory. with an interpretation in terms of questions and answers
Grünwald, P. D. and Vitányi, P. M · 2003
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Bishop, C. M. and Nasrabadi, N. M · 2006
Earlier work this paper cites.
The Minimum Description Length Principle
Grünwald, P. D · 2007
Earlier work this paper cites.
An Introduction to Kolmogorov Complexity and Its Applications
Li, M. and Vitányi, P · 2008
Earlier work this paper cites.
Measuring information transfer in neural networks
Zhang, X., Li, X., Dou, D., and Wu, J · 2009
Earlier work this paper cites.
A complete theory of everything (will be subjective)
Hutter, M · 2010
Earlier work this paper cites.
A philosophical treatise of universal induction
Rathmanner, S. and Hutter, M · 2011
Earlier work this paper cites.
The Beginning of Infinity: Explanations That Transform the World
Deutsch, D · 2012
Earlier work this paper cites.
Intelligence as inference or forcing occam on the world
Sunehag, P. and Hutter, M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Auto-encoders: reconstruction versus compression, 2015
Ollivier, Y · 2015
Earlier work this paper cites.
Compress and control
Veness, J., Bellemare, M. G., Hutter, M., Chua, A., and Desjardins, G · 2015
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T · 2016
Cited alongside, same era.
The description length of deep learning models
Blier, L. and Ollivier, Y · 2018
Cited alongside, same era.
Deepzip: Lossless data compression using recurrent neural networks
Goyal, M., Tatwawadi, K., Chandak, S., and Ochoa, I · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Meta-trained agents implement bayes-optimal agents
Mikulik, V., Delétang, G., McGrath, T., Genewein, T., Martic, M., Legg, S., and Ortega, P · 2020
Cited alongside, same era.
Uncovering mesa-optimization algorithms in transformers
Oswald, J. V., Niklasson, E., Schlegel, M., Kobayashi, S., Zucchet, N., Scherrer, N., Miller, N., Sandler, M., y Arcas, B. A., Vladymyrov, M., Pascanu, R., and Sacramento, J · 2023
Later among the works it cites.
Llmzip: Lossless text compression using large language models, 2023
Valmeekam, C. S. K., Narayanan, K., Kalathil, D., Chamberland, J.-F., and Shakkottai, S · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Large language models are latent variable models: Explaining and finding good demonstrations for in-context learning
Wang, X., Zhu, W., Saxon, M., Steyvers, M., and Wang, W. Y · 2023
Later among the works it cites.
Transformers for supervised online continual learning
Bornschein, J., Li, Y., and Rannen-Triki, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data distributional properties drive emergent in-context learning in transformers
Chan, S., Santoro, A., Lampinen, A., Wang, J., Singh, A., Richemond, P., McClelland, J., and Hill, F · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P., and Valiant, G · 2022
Cited alongside, same era.
General-purpose in-context learning by meta-learning transformers
Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L · 2022
Cited alongside, same era.
Transformers can do bayesian inference
Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Cited alongside, same era.
Closest in time.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Dao, T. and Gu, A · 2024
Closest in time.
Language modeling is compression
Deletang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J · 2024
Closest in time.
CausalLM is not optimal for in-context learning
Ding, N., Levinboim, T., Wu, J., Goodman, S., and Soricut, R · 2024
Closest in time.
Position: The no free lunch theorem, kolmogorov complexity, and the role of inductive biases in machine learning
Goldblum, M., Finzi, M., Rowan, K., and Wilson, A. G · 2024
Closest in time.
Learning universal predictors
Grau-Moya, J., Genewein, T., Hutter, M., Orseau, L., Deletang, G., Catt, E., Ruoss, A., Wenliang, L. K., Mattern, C., Aitchison, M., and Veness, J · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Closest in time.
The broader spectrum of in-context learning
Lampinen, A. K., Chan, S. C., Singh, A. K., and Shanahan, M · 2024
Closest in time.
Structured state space models for in-context reinforcement learning
Lu, C., Schroecker, Y., Gu, A., Parisotto, E., Foerster, J., Singh, S., and Behbahani, F · 2024
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Mahankali, A. V., Hashimoto, T., and Ma, T · 2024
Closest in time.
In-context learning through the bayesian prism
Panwar, M., Ahuja, K., and Goyal, N · 2024
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Raventós, A., Paul, M., Chen, F., and Ganguli, S · 2024
Closest in time.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L · 2024
Closest in time.
Does learning the right latent variables necessarily improve in-context learning?
Mittal, S., Elmoznino, E., Gagnon, L., Bhardwaj, S., Sridhar, D., and Lajoie, G · 2025
Closest in time.
Deep neural networks have an inbuilt Occam’s razor
Mingard, C., Rees, H., Valle-Pérez, G., and Louis, A. A · 2041
Closest in time.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2074
Closest in time.