Fetching the paper…
Reading the bibliography…
Transformers have a remarkable ability to learn and execute tasks based on examples provided within the input itself, without explicit prior training.
A. S. Householder, Unitary triangularization of a nonsymmetric matrix, J. ACM 5
1958
Earlier work this paper cites.
V. A. Marchenko and L. A. Pastur, Distribution of eigenvalues for some sets of random matrices, Matematicheskii Sbornik 114
1967
Earlier work this paper cites.
T. L. H. Watkin, A. Rau, and M. Biehl, The statistical mechanics of learning a rule, Rev. Mod. Phys. 65
1993
Earlier work this paper cites.
L. N. Trefethen and D. Bau, III, Numerical Linear Algebra (Society for Industrial and Applied Mathematics, Philadelphia, PA, 1997) https://epubs.siam.org/doi/pdf/10.1137/1.9780898719574
1997
Earlier work this paper cites.
A. Engel and C. van den Broeck, Statistical Mechanics of Learning (Cambridge University Press, 2001)
2001
Earlier work this paper cites.
2006
Earlier work this paper cites.
A. Tsybakov, Introduction to Nonparametric Estimation , Springer Series in Statistics (Springer New York, 2008)
2008
Earlier work this paper cites.
2010
Earlier work this paper cites.
Z. Bai and J. W. Silverstein, Spectral analysis of large dimensional random matrices , Vol. 20 (Springer, 2010)
2010
Earlier work this paper cites.
L. Erdős, A. Knowles, H.-T. Yau, and J. Yin, The local semicircle law for a general class of random matrices, Electronic Journal of Probability 18
2013
Earlier work this paper cites.
D. Kingma and J. Ba, Adam: A method for stochastic optimization, in International Conference on Learning Representations (ICLR) (San Diega, CA, USA, 2015)
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30
2017
Earlier work this paper cites.
L. Erdős and H.-T. Yau, A dynamical approach to random matrix theory , Vol. 28 (American Mathematical Soc., 2017)
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. , Improving language understanding by generative pre-training (2018)
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. , Language models are unsupervised multitask learners, OpenAI blog 1
2019
Earlier work this paper cites.
M. Belkin, A. Rakhlin, and A. B. Tsybakov, Does data interpolation contradict statistical optimality?, in Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics , Proceedings of Machine Learning Research, Vol. 89, edited by K. Chaudhuri and M. Sugiyama (PMLR, 2019) pp. 1611–1619
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. , Language models are few-shot learners, Advances in neural information processing systems 33
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, Transformers are RNNs: Fast autoregressive transformers with linear attention, in International conference on machine learning (PMLR, 2020) pp. 5156–5165
2020
Earlier work this paper cites.
F. Gerace, B. Loureiro, F. Krzakala, M. Mézard, and L. Zdeborová, Generalisation error in learning with random features and the hidden manifold model, in International Conference on Machine Learning (PMLR, 2020) pp. 3452–3462
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, Efficient attention: Attention with linear complexities, in Proceedings of the IEEE/CVF winter conference on applications of computer vision (2021) pp. 3531–3539
2021
Cited alongside, same era.
B. Loureiro, C. Gerbelot, H. Cui, S. Goldt, F. Krzakala, M. Mezard, and L. Zdeborová, Learning curves of generic features maps for realistic datasets with a teacher-student model, Advances in Neural Information Processing Systems 34
2021
Cited alongside, same era.
A. Bietti, V. Cabannes, D. Bouchacourt, H. Jegou, and L. Bottou, Birth of a transformer: A memory viewpoint, in Advances in Neural Information Processing Systems , Vol. 36, edited by A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Curran Associates, Inc., 2023) pp. 1560–1588
2023
Later among the works it cites.
A. Raventós, M. Paul, F. Chen, and S. Ganguli, Pretraining task diversity and the emergence of non-bayesian in-context learning for regression, in Advances in Neural Information Processing Systems , Vol. 36, edited by A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Curran Associates, Inc., 2023) pp. 14228–14246
2023
Later among the works it cites.
Y. Bai, F. Chen, H. Wang, C. Xiong, and S. Mei, Transformers as statisticians: Provable in-context learning with in-context algorithm selection, in Advances in Neural Information Processing Systems , Vol. 36, edited by A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Curran Associates, Inc., 2023) pp. 57125–57211
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. M. Lu, Householder dice: A matrix-free algorithm for simulating dynamics on Gaussian and random orthogonal ensembles, IEEE Transactions on Information Theory 67
2021
Cited alongside, same era.
D. Ganguli, D. Hernandez, L. Lovitt, A. Askell, Y. Bai, A. Chen, T. Conerly, N. Dassarma, D. Drain, N. Elhage, et al. , Predictability and surprise in large generative models, in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (2022) pp. 1747–1764
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, Emergent abilities of large language models, Transactions on Machine Learning Research (2022)
2022
Cited alongside, same era.
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah, In-context learning and induction heads, Transformer Circuits Thread (2022)
2022
Cited alongside, same era.
B. Barak, B. Edelman, S. Goel, S. Kakade, E. Malach, and C. Zhang, Hidden progress in deep learning: Sgd learns parities near the computational limit, in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022) pp. 21750–21764
2022
Cited alongside, same era.
2022
Cited alongside, same era.
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation, The Annals of Statistics 50
2022
Cited alongside, same era.
2023
Later among the works it cites.
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou, What learning algorithm is in-context learning? investigations with linear models, in The Eleventh International Conference on Learning Representations (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov, Transformers learn in-context by gradient descent, in Proceedings of the 40th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 202, edited by A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (PMLR, 2023) pp. 35151–35174
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Muennighoff, A. M. Rush, B. Barak, T. L. Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. Raffel, Scaling data-constrained language models, in Thirty-seventh Conference on Neural Information Processing Systems (2023)
2023
Later among the works it cites.
G. Reddy, The mechanistic basis of data dependence and abrupt learning in an in-context classification task, in The Twelfth International Conference on Learning Representations (2024)
2024
Closest in time.
J. Wu, D. Zou, Z. Chen, V. Braverman, Q. Gu, and P. Bartlett, How many pretraining tasks are needed for in-context learning of linear regression?, in The Twelfth International Conference on Learning Representations (2024)
2024
Closest in time.
P. Chandra, T. K. Sinha, K. Ahuja, A. Garg, and N. Goyal, Towards analyzing self-attention via linear neural network (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. L. Tong and C. Pehlevan, MLPs learn in-context on regression and classification tasks, in The Thirteenth International Conference on Learning Representations (2025)
2025
Closest in time.