Fetching the paper…
Reading the bibliography…
The emergence of In-Context Learning (ICL) in LLMs remains a remarkable phenomenon that is partially understood.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N · 2018
Earlier work this paper cites.
The CommitmentBank: Investigating projection in naturally occurring discourse
De Marneffe, M.-C., Simons, M., and Tonhauser, J · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
True few-shot learning with language models
Perez, E., Kiela, D., and Cho, K · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2022
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in transformers
Chan, S., Santoro, A., Lampinen, A., Wang, J., Singh, A., Richemond, P., McClelland, J., and Hill, F · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
A theory of emergent in-context learning as implicit structure induction
Hahn, M. and Goyal, N · 2023
Closest in time.
Understanding in-context learning via supportive pretraining data
Han, X., Simig, D., Mihaylov, T., Tsvetkov, Y., Celikyilmaz, A., and Wang, T · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Li, S., Song, Z., Xia, Y., Yu, T., and Zhou, T · 2023
Closest in time.
What in-context learning “learns” in-context: Disentangling task recognition and task learning
Pan, J., Gao, T., Chen, H., and Chen, D · 2023
Closest in time.
Flatness-aware prompt selection improves accuracy and sample efficiency
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2022
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
Merrill, W., Sabharwal, A., and Smith, N. A · 2022
Cited alongside, same era.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Cited alongside, same era.
Reframing instructional prompts to gptk’s language
Mishra, S., Khashabi, D., Baral, C., Choi, Y., and Hajishirzi, H · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot reasoning
Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S · 2022
Cited alongside, same era.
On the effect of pretraining corpora on in-context learning by a large-scale language model
Shin, S., Lee, S. W., Ahn, H., Kim, S., Kim, H. S., Kim, B., Cho, K., Lee, G., Park, W., Ha, J. W., et al · 2022
Cited alongside, same era.
Shen, L., Tan, W., Zheng, B., and Khashabi, D · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Closest in time.
Self-Instruct: Aligning Language Model with Self Generated Instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Closest in time.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Closest in time.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L · 2023
Closest in time.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2024
Closest in time.
Dual operating modes of in-context learning
Lin, Z. and Lee, K · 2024
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Raventós, A., Paul, M., Chen, F., and Ganguli, S · 2024
Closest in time.
The learnability of in-context learning
Wies, N., Levine, Y., and Shashua, A · 2024
Closest in time.