Fetching the paper…
Reading the bibliography…
Pre-trained masked language models have demonstrated remarkable ability as few-shot learners.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Learning and evaluating general linguistic intelligence
Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, et al. 2019 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Exploiting cloze questions for few shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2020a · 2001
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2003
Earlier work this paper cites.
Revisiting few-sample bert fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2006
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze. 2020b · 2009
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2010
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum. 2015 · 2015
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2019 · 2019
Cited alongside, same era.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Medication mention detection in tweets using electra transformers and decision trees
Lung-Hao Lee, Po-Han Chen, Hao-Chuan Kao, Ting-Chun Hung, Po-Lei Lee, and Kuo-Kai Shyu. 2020 · 2020
Cited alongside, same era.
Bertese: Learning to speak to bert
Adi Haviv, Jonathan Berant, and Amir Globerson. 2021 · 2021
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021 · 2021
Later among the works it cites.
Pretraining text encoders with adversarial mixture of training signal generators
Yu Meng, Chenyan Xiong, Payal Bajaj, Paul N Bennett, Jiawei Han, Xia Song, et al. 2021 · 2021
Later among the works it cites.
Noisy channel language model prompting for few-shot text classification
Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021 · 2021
Later among the works it cites.
An empirical comparison of bert, roberta, and electra for fact verification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large biomedical question answering models with albert and electra
Sultan Alrowili and K Shanker. 2021 · 2021
Cited alongside, same era.
Pada: A prompt-based autoregressive approach for adaptation to unseen domains
Eyal Ben-David, Nadav Oved, and Roi Reichart. 2021 · 2021
Cited alongside, same era.
Asr rescoring and confidence estimation with electra
Hayato Futami, Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara. 2021 · 2021
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Cited alongside, same era.
The effect of bert, electra and albert language models on sentiment analysis for turkish product reviews
Zekeriya Anil Guven. 2021 · 2021
Cited alongside, same era.
Warp: Word-level adversarial reprogramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Cited alongside, same era.
Ptr: Prompt tuning with rules for text classification
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Cited alongside, same era.
Muchammad Naseer, Muhamad Asvial, and Riri Fitri Sari. 2021 · 2021
Later among the works it cites.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner. 2021 · 2021
Later among the works it cites.
Efficient passage retrieval with hashing for open-domain question answering
Ikuya Yamada, Akari Asai, and Hannaneh Hajishirzi. 2021 · 2021
Later among the works it cites.
Multi-class grammatical error detection for correction: A tale of two systems
Zheng Yuan, Shiva Taslimipoor, Christopher Davis, and Christopher Bryant. 2021 · 2021
Later among the works it cites.
An emotional classification method of chinese short comment text based on electra
Shunxiang Zhang, Hongbin Yu, and Guangli Zhu. 2021 · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Later among the works it cites.
Corrupted image modeling for self-supervised visual pre-training
Yuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang, and Furu Wei. 2022 · 2022
Closest in time.
Differentiable prompt makes pre-trained language models better few-shot learners
Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022 · 2022
Closest in time.