Fetching the paper…
Reading the bibliography…
The recent surge of large language models (LLMs) highlights their ability to perform in-context learning, i.e., "learning" to perform a task from a few demonstrations in the context without any parameter updates.
Learning multiple visual domains with residual adapters
Rebuffi, S.-A., Bilen, H., and Vedaldi, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Bapna, A. and Firat, O · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Unifiedqa: Crossing format boundaries with a single qa system
Khashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., and Hajishirzi, H · 2020
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, W.-t., Wang, S., and Tang, J · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models. arxiv 2020
Tay, Y., Bahri, D., Metzler, D., Juan, D., Zhao, Z., and Zheng, C · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity, 2020
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D · 2021
Earlier work this paper cites.
Parameter-efficient transfer learning with diff pruning
Guo, D., Rush, A. M., and Kim, Y · 2021
Earlier work this paper cites.
Compacter: Efficient low-rank hypercomplex adapter layers
Henderson, J., Ruder, S., et al · 2021
Earlier work this paper cites.
Surface form competition: Why the highest probability answer isn’t always right
Holtzman, A., West, P., Shwartz, V., Choi, Y., and Zettlemoyer, L · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Cited alongside, same era.
Leveraging passage retrieval with generative models for open domain question answering
Izacard, G. and Grave, É · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Improving in-context few-shot learning via self-supervised training
Chen, M., Du, J., Pasunuru, R., Mihaylov, T., Iyer, S., Stoyanov, V., and Kozareva, Z · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Later among the works it cites.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, X., Ji, K., Fu, Y., Du, Z., Yang, Z., and Tang, J · 2021
Cited alongside, same era.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2021
Cited alongside, same era.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2021
Cited alongside, same era.
cosformer: Rethinking softmax in attention
Qin, Z., Sun, W., Deng, H., Li, D., Wei, Y., Lv, B., Yan, J., Kong, L., and Zhong, Y · 2021
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., and Berant, J · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., et al · 2021
Cited alongside, same era.
Training neural networks with fixed sparse masks
Sung, Y.-L., Nair, V., and Raffel, C. A · 2021
Cited alongside, same era.
Garg, S., Tsipras, D., Liang, P., and Valiant, G · 2022
Later among the works it cites.
Structured prompting: Scaling in-context learning to 1,000 examples
Hao, Y., Sun, Y., Dong, L., Han, Z., Gu, Y., and Wei, F · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Later among the works it cites.
What makes good in-context examples for gpt-3?
Liu, J., Shen, D., Zhang, Y., Dolan, W. B., Carin, L., and Chen, W · 2022
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2022
Later among the works it cites.
Z-icl: Zero-shot in-context learning with pseudo-demonstrations
Lyu, X., Min, S., Beltagy, I., Zettlemoyer, L., and Hajishirzi, H · 2022
Later among the works it cites.
Noisy channel language model prompting for few-shot text classification
Min, S., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Later among the works it cites.
Bidirectional language models are also few-shot learners
Patel, A., Li, B., Rasooli, M. S., Constant, N., Raffel, C., and Callison-Burch, C · 2022
Later among the works it cites.
Reasoning with language model prompting: A survey
Qiao, S., Ou, Y., Zhang, N., Chen, X., Yao, Y., Deng, S., Tan, C., Huang, F., and Chen, H · 2022
Later among the works it cites.
Parallel context windows improve in-context learning of large language models
Ratner, N., Levine, Y., Belinkov, Y., Ram, O., Abend, O., Karpas, E., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Goldberg, Y., and Ravfogel, S · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Transformers as algorithms: Generalization and implicit model selection in in-context learning
Li, Y., Ildiz, M. E., Papailiopoulos, D., and Oymak, S · 2023
Closest in time.