Fetching the paper…
Reading the bibliography…
Neural sequence models based on the transformer architecture have demonstrated remarkable \emph{in-context learning} (ICL) abilities, where they can perform new tasks when prompted with training and test examples, without any parameter update to the model.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Meta-neural networks that learn by learning
D. K. Naik and R. J. Mammone · 1992
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
A model of inductive bias learning
J. Baxter · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. S. Younger, and P. R. Conwell · 2001
Earlier work this paper cites.
On the computational power of transformers and its implications in sequence modeling
S. Bhattamishra, A. Patel, and N. Goyal · 2006
Earlier work this paper cites.
Gradient-based algorithms with applications to signal recovery
A. Beck and M. Teboulle · 2009
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
S. Bhattamishra, K. Ahuja, and N. Goyal · 2009
Earlier work this paper cites.
Fast global convergence rates of gradient methods for high-dimensional statistical recovery
A. Agarwal, S. Negahban, and M. J. Wainwright · 2010
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
S. M. Kakade, V. Kanade, O. Shamir, and A. Kalai · 2011
Earlier work this paper cites.
Random design analysis of ridge regression
D. Hsu, S. M. Kakade, and T. Zhang · 2012
Earlier work this paper cites.
A unified framework for high-dimensional analysis of m-estimators with decomposable regularizers
S. N. Negahban, P. Ravikumar, M. J. Wainwright, and B. Yu · 2012
Earlier work this paper cites.
Learning to learn
S. Thrun and L. Pratt · 2012
Earlier work this paper cites.
On the optimization of a synaptic learning rule
S. Bengio, Y. Bengio, J. Cloutier, and J. Gescei · 2013
Earlier work this paper cites.
Proximal algorithms
N. Parikh, S. Boyd, et al · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
K. Li and J. Malik · 2016
Earlier work this paper cites.
The benefit of multitask representation learning
A. Maurer, M. Pontil, and B. Romera-Paredes · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
F. Bach · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
A simple neural attentive meta-learner
N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel · 2017
Earlier work this paper cites.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
J. Snell, K. Swersky, and R. Zemel · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification
E. Dobriban and S. Wager · 2018
Earlier work this paper cites.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
L. Dong, S. Xu, and B. Xu · 2018
Earlier work this paper cites.
The landscape of empirical risk for nonconvex losses
S. Mei, Y. Bai, and A. Montanari · 2018
Earlier work this paper cites.
Lectures on convex optimization , volume 137
Y. Nesterov · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Cited alongside, same era.
Online meta-learning
C. Finn, A. Rajeswaran, S. Kakade, and S. Levine · 2019
Cited alongside, same era.
Adaptive gradient-based meta-learning methods
M. Khodak, M.-F. F. Balcan, and A. S. Talwalkar · 2019
Cited alongside, same era.
Generalized linear models
P. McCullagh · 2019
Cited alongside, same era.
C. Wei, Y. Chen, and T. Ma · 2021
Later among the works it cites.
Thinking like transformers
G. Weiss, Y. Goldberg, and E. Yahav · 2021
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
S. Yao, B. Peng, C. Papadimitriou, and K. Narasimhan · 2021
Later among the works it cites.
Do transformers really perform badly for graph representation?
C. Ying, T. Cai, S. Luo, S. Zheng, G. Ke, D. He, Y. Shen, and T.-Y. Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Pérez, J. Marinković, and P. Barceló · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
M. J. Wainwright · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
C. Yun, S. Bhojanapalli, A. S. Rawat, S. J. Reddi, and S. Kumar · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Few-shot learning via learning the representation, provably
S. S. Du, W. Hu, S. M. Kakade, J. D. Lee, and Q. Lei · 2020
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh · 2021
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2022
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
S. Chan, A. Santoro, A. Lampinen, J. Wang, A. Singh, P. Richemond, J. McClelland, and F. Hill · 2022
Later among the works it cites.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
D. Dai, Y. Sun, L. Dong, Y. Hao, Z. Sui, and F. Wei · 2022
Later among the works it cites.
A survey for in-context learning
Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, and Z. Sui · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
B. L. Edelman, S. Goel, S. Kakade, and C. Zhang · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
S. Garg, D. Tsipras, P. S. Liang, and G. Valiant · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
S. Jelassi, M. E. Sander, and Y. Li · 2022
Later among the works it cites.
General-purpose in-context learning by meta-learning transformers
L. Kirsch, J. Harrison, J. Sohl-Dickstein, and L. Metz · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
B. Liu, J. T. Ash, S. Goel, A. Krishnamurthy, and C. Zhang · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Later among the works it cites.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot reasoning
Y. Razeghi, R. L. Logan IV, M. Gardner, and S. Singh · 2022
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
J. von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2022
Later among the works it cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Looped transformers as programmable computers
A. Giannou, S. Rajput, J.-y. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos · 2023
Closest in time.
Transformers as algorithms: Generalization and implicit model selection in in-context learning
Y. Li, M. E. Ildiz, D. Papailiopoulos, and S. Oymak · 2023
Closest in time.
OpenAI · 2023
Closest in time.
The effects of pretraining task diversity on in-context learning of ridge regression
A. Raventos, M. Paul, F. Chen, and S. Ganguli · 2023
Closest in time.
A study on relu and softmax in transformer
K. Shen, J. Guo, X. Tan, S. Tang, R. Wang, and J. Bian · 2023
Closest in time.
Larger language models do in-context learning differently
J. Wei, J. Wei, Y. Tay, D. Tran, A. Webson, Y. Lu, X. Chen, H. Liu, D. Huang, D. Zhou, et al · 2023
Closest in time.
Understanding train-validation split in meta-learning with neural networks
X. Zuo, Z. Chen, H. Yao, Y. Cao, and Q. Gu · 2023
Closest in time.