Fetching the paper…
Reading the bibliography…
In-context learning is a remarkable capability of transformers, referring to their ability to adapt to specific tasks based on a short history or context.
Blockwise self-attention for long document understanding
J. Qiu, H. Ma, O. Levy, S. W.-t. Yih, S. Wang, and J. Tang · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. J. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens., 2020
nostalgebraist · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Earlier work this paper cites.
X. Liu, Y. Zheng, Z. Du, M. Ding, Y. Qian, Z. Yang, and J. Tang · 2021
Earlier work this paper cites.
Transformers can do bayesian inference
S. Muller, N. Hollmann, S. P. Arango, J. Grabocka, and F. Hutter · 2021
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2021
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2022
Earlier work this paper cites.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers
D. Dai, Y. Sun, L. Dong, Y. Hao, S. Ma, Z. Sui, and F. Wei · 2022
Earlier work this paper cites.
Analyzing transformers in embedding space
G. Dar, M. Geva, A. Gupta, and J. Berant · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
S. Garg, D. Tsipras, P. Liang, and G. Valiant · 2022
Earlier work this paper cites.
Hyperprompt: Prompt-based task-conditioning of transformers
Y. He, H. S. Zheng, Y. Tay, J. Gupta, Y. Du, V. Aribandi, Z. Zhao, Y. Li, Z. Chen, D. Metzler, H.-T. Cheng, and E. H. Chi · 2022
Earlier work this paper cites.
Editing models with task arithmetic
G. Ilharco, M. T. Ribeiro, M. Wortsman, S. Gururangan, L. Schmidt, H. Hajishirzi, and A. Farhadi · 2022
Earlier work this paper cites.
Ground-truth labels matter: A deeper look into input-label demonstrations
J. Kim, H. J. Kim, H. Cho, H. Jo, S.-W. Lee, S. goo Lee, K. M. Yoo, and T. Kim · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Earlier work this paper cites.
Transformers learn in-context by gradient descent
J. von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2022
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Earlier work this paper cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Y. Bai, F. Chen, H. Wang, C. Xiong, and S. Mei · 2023
Earlier work this paper cites.
Understanding in-context learning in transformers and llms by learning to learn discrete functions
S. Bhattamishra, A. Patel, P. Blunsom, and V. Kanade · 2023
Earlier work this paper cites.
Looped transformers as programmable computers
A. Giannou, S. Rajput, J. yong Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos · 2023
Earlier work this paper cites.
Think before you speak: Training language models with pause tokens
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2023
Earlier work this paper cites.
In-context learning creates task vectors
R. Hendel, M. Geva, and A. Globerson · 2023
Cited alongside, same era.
Chat vector: A simple approach to equip llms with instruction following and model alignment in new languages
S.-C. Huang, P.-Z. Li, Y.-C. Hsu, K.-M. Chen, Y. T. Lin, S.-K. Hsiao, R. T.-H. Tsai, and H. yi Lee · 2023
Cited alongside, same era.
In-context learning learns label relationships but is not conventional learning
J. Kossen, Y. Gal, and T. Rainforth · 2023
Cited alongside, same era.
Supervised pretraining can learn in-context reinforcement learning
J. Lee, A. Xie, A. Pacchiano, Y. Chandak, C. Finn, O. Nachum, and E. Brunskill · 2023
Cited alongside, same era.
Transformers as algorithms: Generalization and stability in in-context learning
Y. Li, M. E. Ildiz, D. Papailiopoulos, and S. Oymak · 2023
Cited alongside, same era.
In-context learning and occam’s razor
E. Elmoznino, T. Marty, T. Kasetty, L. Gagnon, S. Mittal, M. Fathi, D. Sridhar, and G. Lajoie · 2024
Later among the works it cites.
Emergence of abstractions: Concept encoding and decoding mechanism for in-context learning in transformers
S. Han, J. Song, J. Gore, and P. Agrawal · 2024
Later among the works it cites.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms
M. Hanna, S. Pezzelle, and Y. Belinkov · 2024
Later among the works it cites.
A. Hojel, Y. Bai, T. Darrell, A. Globerson, and A. Bar · 2024
Later among the works it cites.
Emotion arithmetic: Emotional speech synthesis via weight space interpolation
P. Kalyan, P. Rao, P. Jyothi, and P. Bhattacharyya · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Lin, Y. Bai, and S. Mei · 2023
Cited alongside, same era.
S. Liu, H. Ye, L. Xing, and J. Y. Zou · 2023
Cited alongside, same era.
Learning to compress prompts with gist tokens
J. Mu, X. L. Li, and N. D. Goodman · 2023
Cited alongside, same era.
What in-context learning "learns" in-context: Disentangling task recognition and task learning
J. Pan, T. Gao, H. Chen, and D. Chen · 2023
Cited alongside, same era.
The effects of pretraining task diversity on in-context learning of ridge regression
A. Raventos, M. Paul, F. Chen, and S. Ganguli · 2023
Cited alongside, same era.
Context compression for auto-regressive transformers with sentinel tokens
S. Ren, Q. Jia, and K. Q. Zhu · 2023
Cited alongside, same era.
Discovering modular solutions that generalize compositionally
S. Schug, S. Kobayashi, Y. Akram, M. Wolczyk, A. Proca, J. von Oswald, R. Pascanu, J. Sacramento, and A. Steger · 2023
Cited alongside, same era.
Later among the works it cites.
Whisper multilingual downstream task tuning using task vectors
J.-H. Kang, J. Lee, M.-H. Lee, and J.-H. Chang · 2024
Later among the works it cites.
When can transformers compositionally generalize in-context?
S. Kobayashi, S. Schug, Y. Akram, F. Redhardt, J. von Oswald, R. Pascanu, G. Lajoie, and J. Sacramento · 2024
Later among the works it cites.
Dual operating modes of in-context learning
Z. Lin and K. Lee · 2024
Later among the works it cites.
Q. Long, Y. Wu, W. Wang, and S. J. Pan · 2024
Later among the works it cites.
Sparser is faster and less is more: Efficient sparse attention for long-range transformers
C. Lou, Z. Jia, Z. Zheng, and K. Tu · 2024
Later among the works it cites.
Task vectors are cross-modal
G. Luo, T. Darrell, and A. Bar · 2024
Later among the works it cites.
Does learning the right latent variables necessarily improve in-context learning?
S. Mittal, E. Elmoznino, L. Gagnon, S. Bhardwaj, D. Sridhar, and G. Lajoie · 2024
Later among the works it cites.
Learnable in-context vector for visual question answering
Y. Peng, C. Hao, X. Yang, J. Peng, X. Hu, and X. Geng · 2024
Later among the works it cites.
Robust concept erasure using task vectors
M. Pham, K. O. Marshall, C. Hegde, and N. Cohen · 2024
Later among the works it cites.
Investigating the effectiveness of hypertuning via gisting
J. Phang · 2024
Later among the works it cites.
Why concepts are (probably) vectors
S. T. Piantadosi, D. C. Muller, J. S. Rule, K. Kaushik, M. I. Gorenstein, E. R. Leib, and E. Sanford · 2024
Later among the works it cites.
Learning task representations from in-context learning
B. Saglam, Z. Yang, D. Kalogerias, and A. Karbasi · 2024
Later among the works it cites.
Where does in-context translation happen in large language models
S. Sia, D. Mueller, and K. Duh · 2024
Later among the works it cites.
A. K. Singh, T. Moskovitz, F. Hill, S. C. Y. Chan, and A. M. Saxe · 2024
Later among the works it cites.
Task arithmetic through the lens of one-shot federated learning
Z. Tao, I. Mason, S. Kulkarni, and X. Boix · 2024
Later among the works it cites.
Scaling up personalized aesthetic assessment via task vector customization
J. Yun and J. Choo · 2024
Later among the works it cites.
Beyond single concept vector: Modeling concept subspace in llms with gaussian distribution
H. Zhao, H. Zhao, B. Shen, A. Payani, F. Yang, and M. Du · 2024
Later among the works it cites.
Vector-icl: In-context learning with continuous vector representations
Y. Zhuang, C. Singh, L. Liu, J. Shang, and J. Gao · 2024
Later among the works it cites.