Fetching the paper…
Reading the bibliography…
In-context learning, a capability that enables a model to learn from input examples on the fly without necessitating weight updates, is a defining characteristic of large language models.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Hoerl, A. E. and Kennard, R. W. (1970) · 1970
Earlier work this paper cites.
Regression and the Moore-Penrose pseudoinverse
Albert, A. E. (1972) · 1972
Earlier work this paper cites.
Probability and measure theory
Ash, R. B. and Doléans-Dade, C. A. (2000) · 2000
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, H. (2000) · 2000
Earlier work this paper cites.
Ridge regression and asymptotic minimax estimation over spheres of growing dimension
Dicker, L. H. (2016) · 2016
Earlier work this paper cites.
Discovering causal signals in images
Lopez-Paz, D., Nishihara, R., Chintala, S., Scholkopf, B., and Bottou, L. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J. (2017) · 2017
Cited alongside, same era.
Asymptotics of ridge (less) regression under general source condition
Richards, D., Mourtada, J., and Rosasco, L. (2021) · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. (2021) · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. (2022) · 2022
Cited alongside, same era.
Transformers generalize differently from information stored in context vs in weights
Chan, S. C., Dasgupta, I., Kim, J., Kumaran, D., Lampinen, A. K., and Hill, F. (2022) · 2022
General-purpose in-context learning by meta-learning transformers
Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L. (2022) · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C. (2022) · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al. (2022) · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F. (2022) · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P., and Valiant, G. (2022) · 2022
Cited alongside, same era.
Li, Y., Ildiz, M. E., Papailiopoulos, D., and Oymak, S. (2023) · 2023
Closest in time.
Gpt-4 technical report
OpenAI (2023) · 2023
Closest in time.
Larger language models do in-context learning differently
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al. (2023) · 2023
Closest in time.