Fetching the paper…
Reading the bibliography…
Demonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks.
On the control of automatic processes: A parallel distributed processing account of the stroop effect
Cohen, J. D., Dunbar, K., and McClelland, J. L. (1990) · 1990
Earlier work this paper cites.
Representing task context: proposals based on a connectionist model of action
Botvinick, M. and Plaut, D. C. (2002) · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Tjong Kim Sang, E. F. and De Meulder, F. (2003) · 2003
Earlier work this paper cites.
Divide and conquer: a defense of functional localizers
Saxe, R., Brett, M., and Kanwisher, N. (2006) · 2006
Earlier work this paper cites.
Evaluating functional localizers: the case of the FFA
Berman, M. G., Park, J., Gonzalez, R., Polk, T. A., Gehrke, A., Knaffla, S., and Jonides, J. (2010) · 2010
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2016) · 2016
Earlier work this paper cites.
Word translation without parallel data
Conneau, A., Lample, G., Ranzato, M., Denoyer, L., and Jégou, H. (2017) · 2017
Earlier work this paper cites.
Distinguishing antonyms and synonyms in a pattern-based neural network
Nguyen, K. A., Schulte im Walde, S., and Vu, N. T. (2017) · 2017
Earlier work this paper cites.
HuggingFace’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. (2019) · 2019
Earlier work this paper cites.
Task representations in neural networks trained to perform many cognitive tasks
Yang, G. R., Joglekar, M. R., Song, H. F., Newsome, W. T., and Wang, X.-J. (2019) · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2020
Earlier work this paper cites.
Transforming task representations to perform novel tasks
Lampinen, A. K. and McClelland, J. L. (2020) · 2020
Earlier work this paper cites.
Rich and lazy learning of task representations in brains and neural networks
Flesch, T., Juechems, K., Dumbalska, T., Saxe, A., and Summerfield, C. (2021) · 2021
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. (2021) · 2021
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. (2022) · 2022
Earlier work this paper cites.
Data distributional properties drive emergent in-context learning in transformers
Chan, S. C. Y., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A. K., Richemond, P. H., Mcclelland, J., and Hill, F. (2022) · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. (2022) · 2022
Earlier work this paper cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. (2022) · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. (2022) · 2022
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H. (2022) · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2022) · 2022
Cited alongside, same era.
Compositional task representations for large language models
Shao, N., Cai, Z., Xu, H., Liao, C., Zheng, Y., and Yang, Z. (2022) · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2022) · 2022
Multimodal task vectors enable many-shot multimodal in-context learning
Huang, B., Mitra, C., Arbelle, A., Karlinsky, L., Darrell, T., and Herzig, R. (2024) · 2024
Later among the works it cites.
Gradient-based inference of abstract task representations for generalization in neural networks
Hummos, A., del Río, F., Wang, B. M., Hurtado, J., Calderon, C. B., and Yang, G. R. (2024) · 2024
Later among the works it cites.
The broader spectrum of in-context learning
Lampinen, A. K., Chan, S. C. Y., Singh, A. K., and Shanahan, M. (2024) · 2024
Later among the works it cites.
ReIFE: Re-evaluating instruction-following evaluation
Liu, Y., Shi, K., Fabbri, A. R., Zhao, Y., Wang, P., Wu, C.-S., Joty, S., and Cohan, A. (2024) · 2024
Later among the works it cites.
The llama 3 herd of models
Llama Team, A. I. . M. (2024) · 2024
Later among the works it cites.
Vision-language models create cross-modal task representations
Luo, G., Darrell, T., and Bar, A. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., , and Wei, J. (2022) · 2022
Cited alongside, same era.
From lazy to rich to exclusive task representations in neural networks and neural codes
Farrell, M., Recanatesi, S., and Shea-Brown, E. (2023) · 2023
Cited alongside, same era.
In-context learning creates task vectors
Hendel, R., Geva, M., and Globerson, A. (2023) · 2023
Cited alongside, same era.
Linearity of relation decoding in transformer language models
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D. (2023) · 2023
Cited alongside, same era.
What in-context learning “learns” in-context: Disentangling task recognition and task learning
Pan, J., Gao, T., Chen, H., and Chen, D. (2023) · 2023
Cited alongside, same era.
Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. (2023) · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Cited alongside, same era.
Later among the works it cites.
HREF: Human response-guided evaluation of instruction following in language models
Lyu, X., Wang, Y., Hajishirzi, H., and Dasigi, P. (2024) · 2024
Later among the works it cites.
2 OLMo 2 furious
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., Lambert, N., Schwenk, D., Tafjord, O., Anderson, T., Atkinson, D., Brahman, F., Clark, C., Dasigi, P., Dziri, N., Guerquin, M., Ivison, H., Koh, P. W., Liu, J., Malik, S., Merrill, W., Miranda, L. J. V., Morrison, J., Murray, T., Nam, C., Pyatkin, V., Rangapur, A., Schmitz, M., Skjonsberg, S., Wadden, D., Wilhelm, C., Wilson, M., Zettlemoyer, L., Farhadi, A., Smith, N. A., and Hajishirzi, H. (2024) · 2024
Later among the works it cites.
Competition dynamics shape algorithmic phases of in-context learning
Park, C. F., Lubana, E. S., Pres, I., and Tanaka, H. (2024) · 2024
Later among the works it cites.
Improving instruction-following in language models through activation steering
Stolfo, A., Balachandran, V., Yousefi, S., Horvitz, E., and Nushi, B. (2024) · 2024
Later among the works it cites.
Function vectors in large language models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. (2024) · 2024
Later among the works it cites.
From language modeling to instruction following: Understanding the behavior shift in LLMs after instruction tuning
Wu, X., Yao, W., Chen, J., Pan, X., Wang, X., Liu, N., and Yu, D. (2024) · 2024
Later among the works it cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2021) · 2025
Closest in time.
Shared global and local geometry of language model embeddings
Lee, A., Weber, M., Viégas, F., and Wattenberg, M. (2025) · 2025
Closest in time.
Just-in-time and distributed task representations in language models
Li, Y., Campbell, D., Chan, S. C. Y., and Lampinen, A. K. (2025) · 2025
Closest in time.
Learning task representations from in-context learning
Saglam, B., Yang, Z., Kalogerias, D., and Karbasi, A. (2025) · 2025
Closest in time.
A single character can make or break your LLM evals
Su, J., Zhang, J., Ullrich, K., Bottou, L., and Ibrahim, M. (2025) · 2025
Closest in time.
Which attention heads matter for in-context learning?
Yin, K. and Steinhardt, J. (2025) · 2025
Closest in time.