Fetching the paper…
Reading the bibliography…
In recent years, pre-trained large language models (LLMs) have demonstrated remarkable efficiency in achieving an inference-time few-shot learning capability known as in-context learning.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
A probabilistic theory of pattern recognition
L. Devroye, L. Györfi, and G. Lugosi · 1996
Earlier work this paper cites.
Latent dirichlet allocation
D. M. Blei, A. Y. Ng, and M. I. Jordan · 2003
Earlier work this paper cites.
Semeval-2019 task 3: Emocontext contextual emotion detection in text
A. Chatterjee, K. N. Narahari, M. Joshi, and P. Agrawal · 2005
Earlier work this paper cites.
Visualizing data using t-sne
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala · 2014
Earlier work this paper cites.
Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia
J. Lehmann, R. Isele, M. Jakob, A. Jentzsch, D. Kontokostas, P. N. Mendes, S. Hellmann, M. Morsey, P. Van Kleef, S. Auer, et al · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Earlier work this paper cites.
Neural variational inference for text processing
Y. Miao, L. Yu, and P. Blunsom · 2016
Earlier work this paper cites.
Discovering discrete latent topics with neural variational inference
Y. Miao, E. Grefenstette, and P. Blunsom · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
E. Saravia, H.-C. T. Liu, Y.-H. Huang, J. Wu, and Y.-S. Chen · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Neural network acceptability judgments
A. Warstadt, A. Singh, and S. R. Bowman · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Ethos: an online hate speech detection dataset, 2020
I. Mollas, Z. Chrysopoulou, S. Karlos, and G. Tsoumakas · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Cited alongside, same era.
True few-shot learning with language models
What makes good in-context examples for GPT-3?
J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen · 2022
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp · 2022
Later among the works it cites.
MetaICL: Learning to learn in context
S. Min, M. Lewis, L. Zettlemoyer, and H. Hajishirzi · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Selective annotation makes language models better few-shot learners
H. Su, J. Kasai, C. H. Wu, W. Shi, T. Wang, J. Xin, R. Zhang, M. Ostendorf, L. Zettlemoyer, N. A. Smith, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Perez, D. Kiela, and K. Cho · 2021
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
O. Rubin, J. Herzig, and J. Berant · 2021
Cited alongside, same era.
A mathematical exploration of why language models help solve downstream tasks
N. Saunshi, S. Malladi, and S. Arora · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
B. Wang and A. Komatsuzaki · 2021
Cited alongside, same era.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
C. Wei, S. M. Xie, and T. Ma · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2022
Cited alongside, same era.
H. Bansal, K. Gopalakrishnan, S. Dingliwal, S. Bodapati, K. Kirchhoff, and D. Roth · 2022
Cited alongside, same era.
Transformers learn in-context by gradient descent
J. von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le, and E. Chi · 2022
Later among the works it cites.
J. Chen, L. Chen, and T. Zhou · 2023
Closest in time.
A theory of emergent in-context learning as implicit structure induction
M. Hahn and N. Goyal · 2023
Closest in time.
Understanding in-context learning via supportive pretraining data
X. Han, D. Simig, T. Mihaylov, Y. Tsvetkov, A. Celikyilmaz, and T. Wang · 2023
Closest in time.
A latent space theory for emergent abilities in large language models
H. Jiang · 2023
Closest in time.
Transformers as algorithms: Generalization and implicit model selection in in-context learning
Y. Li, M. E. Ildiz, D. Papailiopoulos, and S. Oymak · 2023
Closest in time.