Fetching the paper…
Reading the bibliography…
Transformer models, notably large language models (LLMs), have the remarkable ability to perform in-context learning (ICL) -- to perform new tasks when prompted with unseen input-output examples without any explicit model training.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
S. Garg, D. Tsipras, P. S. Liang, and G. Valiant · 2022
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection, 2023
Y. Bai, F. Chen, H. Wang, C. Xiong, and S. Mei · 2023
Cited alongside, same era.
Transformers as algorithms: Generalization and stability in in-context learning
Y. Li, M. Emrullah Ildiz, D. Papailiopoulos, and S. Oymak · 2023
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression, 2023
A. Raventós, M. Paul, F. Chen, and S. Ganguli · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…