Fetching the paper…
Reading the bibliography…
Previous theoretical results pertaining to meta-learning on sequences build on contrived assumptions and are somewhat convoluted.
A model of inductive bias learning
Baxter, J · 2000
Earlier work this paper cites.
Learning bounds for support vector machines with learned kernels
Srebro, N. and Ben-David, S · 2006
Earlier work this paper cites.
Transfer bounds for linear feature learning
Maurer, A · 2009
Earlier work this paper cites.
Excess risk bounds for multitask learning with trace norm regularization
Pontil, M. and Maurer, A · 2013
Earlier work this paper cites.
The benefit of multitask representation learning
Maurer, A., Pontil, M., and Romera-Paredes, B · 2016
Earlier work this paper cites.
Incremental learning-to-learn with statistical guarantees
Denevi, G., Ciliberto, C., Stamos, D., and Pontil, M · 2018
Earlier work this paper cites.
Online meta-learning
Finn, C., Rajeswaran, A., Kakade, S., and Levine, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L., Goel, S., Kakade, S. M., and Zhang, C · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Earlier work this paper cites.
What makes good in-context examples for gpt- 3 3 ?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., and Chen, W · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2021
Earlier work this paper cites.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2021
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., and Berant, J · 2021
Cited alongside, same era.
Provable meta-learning of linear representations
Tripuraneni, N., Jin, C., and Jordan, M · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2022
Cited alongside, same era.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F · 2022
Cited alongside, same era.
Metalearning with very few samples per task
Aliakbarpour, M., Bairaktari, K., Brown, G., Smith, A., and Ullman, J · 2023
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection, 2023
Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S · 2023
Later among the works it cites.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Later among the works it cites.
In-context convergence of transformers
Huang, Y., Cheng, Y., and Liang, Y · 2023
Later among the works it cites.
An information-theoretic framework for supervised learning, 2023
Jeon, H. J., Zhu, Y., and Van Roy, B · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey for in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M., and Li, Y · 2022
Cited alongside, same era.
General-purpose in-context learning by meta-learning transformers
Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L · 2022
Cited alongside, same era.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?, 2022
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference, 2022
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Cited alongside, same era.
Later among the works it cites.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Later among the works it cites.
The effects of pretraining task diversity on in-context learning of ridge regression
Raventos, A., Paul, M., Chen, F., and Ganguli, S · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
Sanford, C., Hsu, D., and Telgarsky, M · 2023
Later among the works it cites.
Uncovering hidden geometry in transformers via disentangling position and context
Song, J. and Zhong, Y · 2023
Later among the works it cites.
Transformers as support vector machines
Tarzanagh, D. A., Li, Y., Thrampoulidis, C., and Oymak, S · 2023
Later among the works it cites.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Larger language models do in-context learning differently
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al · 2023
Later among the works it cites.