Fetching the paper…
Reading the bibliography…
In-context learning (ICL) is a powerful ability that emerges in transformer models, enabling them to learn from context without weight updates.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge, 2015
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Training pruned neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Training dynamics of in-context learning in linear attention, 2025
Zhang, Y., Singh, A. K., Latham, P. E., and Saxe, A · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories, 2021
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Earlier work this paper cites.
Data distributional properties drive emergent in-context learning in transformers
Chan, S., Santoro, A., Lampinen, A., Wang, J., Singh, A., Richemond, P., McClelland, J., and Hill, F · 2022
Cited alongside, same era.
Toy models of superposition
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Cited alongside, same era.
The neural race reduction: Dynamics of abstraction in gated networks, 2022
Saxe, A. M., Sodhani, S., and Lewallen, S · 2022
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Anand, S., Lepori, M. A., Merullo, J., and Pavlick, E · 2024
Later among the works it cites.
Toward Understanding In-context vs. In-weight Learning, October 2024
Chan, B., Chen, X., György, A., and Schuurmans, D · 2024
Later among the works it cites.
He, T., Doshi, D., Das, A., and Gromov, A · 2024
Later among the works it cites.
The broader spectrum of in-context learning, 2024
Lampinen, A. K., Chan, S. C. Y., Singh, A. K., and Shanahan, M · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gaussian error linear units (gelus), 2023
Hendrycks, D. and Gimpel, K · 2023
Cited alongside, same era.
Locating and editing factual associations in gpt, 2023
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability, 2023
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Cited alongside, same era.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task, 2023
Reddy, G · 2023
Cited alongside, same era.
The Transient Nature of Emergent In-Context Learning in Transformers, December 2023
Singh, A. K., Chan, S. C. Y., Moskovitz, T., Grant, E., Saxe, A. M., and Hill, F · 2023
Cited alongside, same era.
Larger language models do in-context learning differently
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al · 2023
Cited alongside, same era.
Lin, Z. and Lee, K · 2024
Later among the works it cites.
Nguyen, A. and Reddy, G · 2024
Later among the works it cites.
Competition Dynamics Shape Algorithmic Phases of In-Context Learning, December 2024
Park, C. F., Lubana, E. S., Pres, I., and Tanaka, H · 2024
Later among the works it cites.
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation, 2024
Singh, A. K., Moskovitz, T., Hill, F., Chan, S. C. Y., and Saxe, A. M · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Wang, X., Tang, X., Zhao, W. X., and Wen, J.-R · 2024
Later among the works it cites.
Which attention heads matter for in-context learning?, 2025
Yin, K. and Steinhardt, J · 2025
Closest in time.