Fetching the paper…
Reading the bibliography…
In-context learning (ICL) exhibits dual operating modes: task learning, i.e., acquiring a new skill from in-context samples, and task retrieval, i.e., locating and activating a relevant pretrained skill.
A tutorial on hidden markov models and selected applications in speech recognition
Rabiner, L. R · 1989
Earlier work this paper cites.
Factorial hidden markov models
Ghahramani, Z. and Jordan, M · 1995
Earlier work this paper cites.
Detection, estimation, and modulation theory, Part I: Detection, estimation, and linear modulation theory
Van Trees, H. L · 2004
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marelli, M., Menini, S., Baroni, M., Bentivogli, L., Bernardi, R., and Zamparelli, R · 2014
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L · 2017
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Vershynin, R · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Tweeteval: Unified benchmark and comparative evaluation for tweet classification
Barbieri, F., Camacho-Collados, J., Anke, L. E., and Neves, L · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Investigating societal biases in a poetry composition system
Sheng, E. and Uthus, D · 2020
Earlier work this paper cites.
What can Transformers learn in-context? A case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Razeghi, Y., IV, R. L. L., Gardner, M., and Singh, S · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit Bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Cited alongside, same era.
What learning algorithm is in-context learning? Investigations with linear models
Z-ICL: Zero-shot in-context learning with pseudo-demonstrations
Lyu, X., Min, S., Beltagy, I., Zettlemoyer, L., and Hajishirzi, H · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
What in-context learning “learns” in-context: Disentangling task recognition and task learning
Pan, J., Gao, T., Chen, H., and Chen, D · 2023
Later among the works it cites.
The effects of pretraining task diversity on in-context learning of ridge regression
Raventos, A., Paul, M., Chen, F., and Ganguli, S · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S · 2023
Cited alongside, same era.
Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F · 2023
Cited alongside, same era.
Looped Transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Cited alongside, same era.
In-context learning of large language models explained as kernel regression
Han, C., Wang, Z., Zhao, H., and Ji, H · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Transformers as algorithms: Generalization and stability in in-context learning
Li, Y., Ildiz, M. E., Papailiopoulos, D., and Oymak, S · 2023
Cited alongside, same era.
Tsigler, A. and Bartlett, P. L · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L · 2023
Later among the works it cites.
An information-theoretic analysis of in-context learning
Jeon, H. J., Lee, J. D., Lei, Q., and Van Roy, B · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Mahankali, A., Hashimoto, T. B., and Ma, T · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
Wu, J., Zou, D., Chen, Z., Braverman, V., Gu, Q., and Bartlett, P. L · 2024
Closest in time.