Fetching the paper…
Reading the bibliography…
Through in-context learning (ICL), large-scale language models are effective few-shot learners without additional model fine-tuning.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T.; Kaplan, J.; Katz, M.; Chen, M.; Hesse, C.; Jackson, J.; Jun, H.; Brown, T. B.; Dhariwal, P.; Gray, S.; et al. 2020 · 2010
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Riedel, S.; Lewis, P. S. H.; Bakhtin, A.; Wu, Y.; and Miller, A. H. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2020
Earlier work this paper cites.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Min, S.; Lyu, X.; Holtzman, A.; Artetxe, M.; Lewis, M.; Hajishirzi, H.; and Zettlemoyer, L. 2022b · 2020
Earlier work this paper cites.
Unsupervised Commonsense Question Answering with Self-Talk
Shwartz, V.; West, P.; Le Bras, R.; Bhagavatula, C.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
Ground-Truth Labels Matter: A Deeper Look into Input-Label Demonstrations
Yoo, K. M.; Kim, J.; Kim, H. J.; Cho, H.; Jo, H.; Lee, S.-W.; Lee, S.-g.; and Kim, T. 2022 · 2020
Earlier work this paper cites.
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Aghajanyan, A.; Gupta, S.; and Zettlemoyer, L. 2021 · 2021
Earlier work this paper cites.
SimCSE: Simple Contrastive Learning of Sentence Embeddings
Gao, T.; Yao, X.; and Chen, D. 2021 · 2021
Cited alongside, same era.
Surface Form Competition: Why the Highest Probability Answer Isn’t Always Right
Holtzman, A.; West, P.; Shwartz, V.; Choi, Y.; and Zettlemoyer, L. 2021 · 2021
Cited alongside, same era.
Liu, X.; Zheng, Y.; Du, Z.; Ding, M.; Qian, Y.; Yang, Z.; and Tang, J. 2021 · 2021
Cited alongside, same era.
Exploring low-dimensional intrinsic task subspace via prompt tuning
Qin, Y.; Wang, X.; Su, Y.; Lin, Y.; Ding, N.; Liu, Z.; Li, J.; Hou, L.; Li, P.; Sun, M.; et al. 2021 · 2021
Cited alongside, same era.
Prompt programming for large language models: Beyond the few-shot paradigm
Reynolds, L.; and McDonell, K. 2021 · 2021
Cited alongside, same era.
Black-box prompt learning for pre-trained language models
Diao, S.; Li, X.; Lin, Y.; Huang, Z.; and Zhang, T. 2022 · 2022
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, W.; Zoph, B.; and Shazeer, N. 2022 · 2022
Closest in time.
Training Compute-Optimal Large Language Models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; Casas, D. d. L.; Hendricks, L. A.; Welbl, J.; Clark, A.; et al. 2022 · 2022
Closest in time.
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Lu, Y.; Bartolo, M.; Moore, A.; Riedel, S.; and Stenetorp, P. 2022 · 2022
Closest in time.
Synchromesh: Reliable Code Generation from Pre-trained Language Models
Poesia, G.; Polozov, A.; Le, V.; Tiwari, A.; Soares, G.; Meek, C.; and Gulwani, S. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Schick, T.; and Schütze, H. 2021b · 2021
Cited alongside, same era.
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Sun, Y.; Wang, S.; Feng, S.; Ding, S.; Pang, C.; Shang, J.; Liu, J.; Chen, X.; Zhao, Y.; Lu, Y.; et al. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Calibrate Before Use: Improving Few-shot Performance of Language Models
Zhao, Z.; Wallace, E.; Feng, S.; Klein, D.; and Singh, S. 2021 · 2021
Cited alongside, same era.
PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen Domains
Ben-David, E.; Oved, N.; and Reichart, R. 2022 · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Cited alongside, same era.
What Makes Good In-Context Examples for GPT-3?
Liu, J.; Shen, D.; Zhang, Y.; Dolan, B.; Carin, L.; and Chen, W. 2022a
Cited in the paper.
Closest in time.
Impact of pretraining term frequencies on few-shot reasoning
Razeghi, Y.; Logan IV, R. L.; Gardner, M.; and Singh, S. 2022 · 2022
Closest in time.
Learning To Retrieve Prompts for In-Context Learning
Rubin, O.; Herzig, J.; and Berant, J. 2022 · 2022
Closest in time.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Closest in time.
Black-Box Tuning for Language-Model-as-a-Service
Sun, T.; Shao, Y.; Qian, H.; Huang, X.; and Qiu, X. 2022 · 2022
Closest in time.
An Explanation of In-context Learning as Implicit Bayesian Inference
Xie, S. M.; Raghunathan, A.; Liang, P.; and Ma, T. 2022 · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Closest in time.