Fetching the paper…
Reading the bibliography…
At present, the mechanisms of in-context learning in Transformers are not well understood and remain mostly an intuition.
Risks from learned optimization in advanced machine learning systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S · 1906
Earlier work this paper cites.
Adaptive switching circuits
Widrow, B. and Hoff, M. E · 1960
Earlier work this paper cites.
On estimating regression
Nadaraya, E. A · 1964
Earlier work this paper cites.
Smooth regression analysis
Watson, G. S · 1964
Earlier work this paper cites.
Using fast weights to deblur old memories
Hinton, G. E. and Plaut, D. C · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Y., Bengio, S., and Cloutier, J · 1990
Earlier work this paper cites.
The evolution of learning: an experiment in genetic connectionism
Chalmers, D. J · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Schmidhuber, J · 1992
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N · 2016
Earlier work this paper cites.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Optnet: Differentiable optimization as a layer in neural networks
Amos, B. and Kolter, J. Z · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Meta-SGD: Learning to learn quickly for few shot learning
Li, Z., Zhou, F., Chen, F., and Li, H · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Finn, C. and Levine, S · 2018
Cited alongside, same era.
Gradient-based meta-learning with learned layerwise metric and subspace
Lee, Y. and Choi, S · 2018
Cited alongside, same era.
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Cited alongside, same era.
Meta-learning with differentiable closed-form solvers
Bertinetto, L., Henriques, J. F., Torr, P. H. S., and Vedaldi, A · 2019
Cited alongside, same era.
Meta-learning probabilistic inference for prediction
Gordon, J., Bronskill, J., Bauer, M., Nowozin, S., and Turner, R · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2021
Later among the works it cites.
Deep declarative networks
Gould, S., Hartley, R., and Campbell, D. J · 2021
Later among the works it cites.
Going beyond linear transformers with recurrent fast weight programmers
Irie, K., Schlag, I., Csordás, R., and Schmidhuber, J · 2021
Later among the works it cites.
Meta learning backpropagation and improving it
Kirsch, L. and Schmidhuber, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta-learning with differentiable convex optimization
Lee, K., Maji, S., Ravichandran, A., and Soatto, S · 2019
Cited alongside, same era.
Meta-curvature
Park, E. and Oliva, J. B · 2019
Cited alongside, same era.
Meta-learning with latent embedding optimization
Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., and Hadsell, R · 2019
Cited alongside, same era.
Graph transformer networks
Yun, S., Jeong, M., Kim, R., Kang, J., and Kim, H. J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Later among the works it cites.
Learning where to learn: Gradient sparsity in meta and continual learning
von Oswald, J., Zhao, D., Kobayashi, S., Schug, S., Caccia, M., Zucchet, N., and Sacramento, J · 2021
Later among the works it cites.
Zhang, A., Lipton, Z. C., Li, M., and Smola, A. J · 2021
Later among the works it cites.
Random initialisations performing above chance and how to find them
Benzing, F., Schug, S., Meier, R., von Oswald, J., Akram, Y., Zucchet, N., Aitchison, L., and Steger, A · 2022
Closest in time.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P., and Valiant, G · 2022
Closest in time.
General-purpose in-context learning by meta-learning transformers
Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L · 2022
Closest in time.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Closest in time.
HyperTransformer: Model generation for supervised and semi-supervised few-shot learning
Zhmoginov, A., Sandler, M., and Vladymyrov, M · 2022
Closest in time.
Beyond backpropagation: bilevel optimization through implicit differentiation and equilibrium propagation
Zucchet, N. and Sacramento, J · 2022
Closest in time.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Closest in time.
Why can GPT learn in-context? language models implicitly perform gradient descent as meta-optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F · 2023
Closest in time.