Fetching the paper…
Reading the bibliography…
In this paper, we conduct a comprehensive study of In-Context Learning (ICL) by addressing several open questions: (a) What type of ICL estimator is learned by large language models? (b) What is a proper performance metric for ICL and what is the error rate? (c) How does the transformer architecture enable ICL? To answer these questions, we adopt a Bayesian view and formulate ICL as a problem of predicting the response corresponding to the current covariate, given a number of examples drawn from a latent variable model.
Tinybert: Distilling bert for natural language understanding
Jiao, X · 1909
Earlier work this paper cites.
Generalization bounds for convolutional neural networks
Lin, S · 1910
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C · 1912
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K · 1991
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Anthony, M · 1999
Earlier work this paper cites.
Bayesian model selection and model averaging
Wasserman, L · 2000
Earlier work this paper cites.
Rethinking positional encoding in language pre-training
Ke, G · 2006
Earlier work this paper cites.
Tensor programs II: Neural tangent kernel for any architecture
Yang, G · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A · 2007
Earlier work this paper cites.
A mathematical theory of attention
Vuckovic, J · 2007
Earlier work this paper cites.
Hilbert space embeddings of conditional distributions with applications to dynamical systems
Song, L · 2009
Earlier work this paper cites.
A pac-bayesian approach to generalization bounds for graph neural networks
Liao, R · 2012
Earlier work this paper cites.
Nonparametric bayesian inference with kernel mean embedding
Fukumizu, K · 2015
Earlier work this paper cites.
Concentration inequalities for markov chains by marton couplings and spectral methods
Paulin, D · 2015
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L · 2017
Earlier work this paper cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Yarotsky, D · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M · 2017
Earlier work this paper cites.
Mutual information neural estimation
Belghazi, M. I · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Dong, L · 2019
Cited alongside, same era.
Information theory and statistics
Duchi, J. C · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank MDPs
Agarwal, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T · 2020
Cited alongside, same era.
What can transformers learn in-context? A case study of simple function classes
Garg, S · 2022
Later among the works it cites.
Instruction induction: From few examples to natural language task descriptions
Honovich, O · 2022
Later among the works it cites.
OPT-IML: Scaling language model instruction meta learning through the lens of generalization
Iyer, S · 2022
Later among the works it cites.
Kim, H. J · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Infinite attention: NNGP and NTK for deep attention networks
Hron, J · 2020
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L · 2021
Cited alongside, same era.
Deep neural network approximation theory
Elbrächter, D · 2021
Cited alongside, same era.
Norm-based generalisation bounds for deep multi-class convolutional neural networks
Ledent, A · 2021
Cited alongside, same era.
What makes good in-context examples for gpt- 3 3 ?
Liu, J · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y · 2021
Cited alongside, same era.
Kojima, T · 2022
Later among the works it cites.
A kernel-based view of language model fine-tuning
Malladi, S · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Noci, L · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
Wang, Y · 2022
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Bai, Y · 2023
Closest in time.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Feng, G · 2023
Closest in time.
Exploring the space of key-value-query models with intention
Garnelo, M · 2023
Closest in time.
A theory of emergent in-context learning as implicit structure induction
Hahn, M · 2023
Closest in time.
A latent space theory for emergent abilities in large language models
Jiang, H · 2023
Closest in time.
Transformers as algorithms: Generalization and stability in in-context learning
Li, Y · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P · 2023
Closest in time.
Wang, X · 2023
Closest in time.
The learnability of in-context learning
Wies, N · 2023
Closest in time.